Microsoft's ThinkingBox grades AI agents on what they leave behind
Microsoft and Hugging Face have published ThinkingBox, a benchmark that scores an AI agent on the state it leaves in a database rather than the answer it types back.

Microsoft and Hugging Face have published a benchmark that scores AI agents on the mess they leave behind rather than the messages they write. ThinkingBox runs an agent through 507 stateful business workflows twenty times each, then inspects the records and side effects sitting in the backend instead of the tool calls it made on the way.
Why the transcript is not the evidence
The post opens with a worked example. An agent is handed a customer whose kitchen appliance has been stuck in a courier “exception” for fifteen days. It makes nine well-formed tool calls: it pulls the order, checks the tracking, reads the customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline and correctly concludes the account does not qualify for compensation. Then it closes the ticket as resolved and asks whether there is anything else it can help with.
Two things are wrong. The carrier exception is still open, so the required end state was on hold rather than finished, and the customer never got an answer to the question she actually asked. A grader reading tool calls sees nine clean ones; a grader checking whether the agent wrote to the database sees a write. Only the database disagrees.
ThinkingBox is the sandbox and ThinkingBox-Bench is the dataset. Each task ships a starting backend state, a user goal, the MCP tools on offer, a domain policy and executable checks. A simulated user holds private context, such as a booking reference or a date of birth, and releases it only when asked. Every attempt gets its own MCP session with freshly initialised state, so two runs of one task never share a database row or cached tool output. A side-effect extractor then derives what actually changed, and deterministic judges compare that with the required end state, accepting any route that gets there and rejecting wrong, missing or extra effects. Four hundred and seventy-seven of the 507 tasks can be graded this way, with the remainder decided by a narrow yes-or-no rubric question.
Who leads, and by how much
Claude Opus 5.5 tops the table at 67.16% pass@1, about two-thirds of a point clear of Claude Opus 5. Kimi-K3 is the strongest open-weights model and lands within a point of GPT-6-Astra. Domain turns out to matter as much as model choice: Claude Opus 4.6 scores 68.62% on retail workflows and 8.30% on auto insurance. Across the models in the paper's second table, retail averages 59.52% pass@1 against 33.83% for auto insurance.
Twenty attempts, and the number that survives
One good run proves a model can do the work once. It does not prove it will do it again, so every task is repeated twenty times and the question becomes how much of that first score survives. GPT-6 Astra retains 78% of its single-attempt rate, Claude Opus 5.5 and Claude Opus 5 each retain 71%, while GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro keep roughly 8%.
Breadth and consistency pull against each other. Kimi-K3 solves 93.89% of the benchmark at least once, or 476 of the 507 tasks, with only 31 defeating it outright, and it leads retail at 82.24%. It also finishes just 68 tasks, 13.41%, on all twenty attempts. Claude Opus 5 inverts the shape: fewer tasks solved at least once at 79.09%, with 106 defeats, but 47.53% of the benchmark completed on every single attempt.
What the authors suggest
The advice in the post is to treat the twenty-out-of-twenty rate as a design input rather than a verdict. Check the terminal state before committing a change instead of trusting the model's own summary of it, classify tool and system errors so that retries target the recoverable ones, trim the tool surface down to what the workflow actually needs, and require human approval for changes that cannot be cheaply reversed. The authors are candid that they have not measured how much those fixes lift scores, which is exactly what the environment now makes testable. The benchmark runs through OpenEnv and is published on Hugging Face.
Our opinion
The interesting thing about ThinkingBox is not the leaderboard, it is where the authors put the finish line. Graders that read tool calls reward an agent for looking busy, and the customer in the opening example received a polite, well-behaved and completely useless answer. Marking the database instead is more expensive to build and far harder to game, which is the whole point.
The consistency column is the one to take into a procurement meeting. Any vendor can show a demo where the agent books the refund correctly, and a single clean run is what most pilot programmes actually measure. Knowing that a strong model keeps roughly 8% of its score across twenty attempts is the difference between buying a capability and buying a lucky afternoon.