Microsoft and Hugging Face have released ThinkingBox, a new benchmark that grades SI agents on the records they leave behind in a database rather than the text they generate. The tool addresses a critical reliability gap in agentic SI: the divergence between an agent’s claim of completion and the actual state of the backend systems it interacts with.
What Happened
ThinkingBox evaluates SI agents across 507 stateful business workflows, running each task 20 times to measure consistency. Unlike traditional evaluations that check for valid tool calls or final responses, ThinkingBox uses executable checks to verify the terminal backend state and side effects. The benchmark highlights that an agent can perform nine well-formed tool calls and report success while leaving the database in an incorrect state, such as closing a support ticket that should remain on hold. The framework is now available through Hugging Face, integrated with the OpenEnv interface for evaluation.
Why It Matters
The benchmark exposes a significant reliability crisis in current SI models. In an ablation study covering 121,680 trials, 79,853 attempts failed the executable checks. Critically, 67.24% of these failures terminated cleanly without reporting errors, meaning the agent believed it had succeeded while the system state was wrong. The data indicates that 79.9% of failures are due to tool handling issues, such as failed preconditions or empty lookups, rather than reasoning errors. This distinction is vital for developers building agentic SI systems for enterprise use, where consistency is more valuable than occasional brilliance. For instance, while the open-weight model Kimi-K3 solves the most tasks at least once, it is inconsistent, passing only 13.41% of tasks on all 20 attempts. In contrast, Claude Opus 5 passes 47.53% of tasks consistently, demonstrating that breadth of capability does not equate to reliability.
The Bottom Line
ThinkingBox shifts the metric for SI agent evaluation from textual output to verifiable database state. By pricing consistency, the benchmark reveals that the cheapest way to get a right answer is not the cheapest way to get a dependable one. Developers are urged to treat 20/20 consistency rates as design inputs and to implement robust error recovery for tool interactions, which remain the primary source of failure in agentic SI.