The Agent Said It Was Done. The Database Disagreed.

2026-10-05 · Hugging Face

The Agent Said It Was Done. The Database Disagreed.

Microsoft ThinkingBox, in collaboration with Hugging Face, introduces a novel benchmark for evaluating AI agents. The core premise is that agents should be graded on the terminal backend state and side effects they leave behind, rather than the sentences they generate. Furthermore, it tests whether an agent can reliably perform the same task twenty times in a row. The benchmark is now available through Hugging Face.

A Tool Call is Not an Outcome

The article illustrates the evaluation gap with a retail customer service case. A customer's $745 kitchen appliance is stuck in a courier exception for 15 days. The AI agent executes nine seemingly perfect tool calls: pulling the order, checking tracking, looking up the profile, searching refund policies, and creating a ticket. However, the agent closes the ticket as "solved" instead of the required "hold" status, and the customer never receives a real answer to their query.

ThinkingBox measures this discrepancy between apparent success and actual backend state. In a common-set ablation covering 121,680 valid trials across 12 LLMs, 79,853 attempts failed executable checks. Among these failures, 67.24% terminated cleanly without reporting any final tool error. However, executable state checks revealed deeper issues:

  • 77.61% contained wrong field values.
  • 43.30% generated unintended extra side effects.
  • 25.36% missed required effects entirely.

A trajectory is merely a claim; the database state is the evidence.

One Success is Not Reliability

An agent that processes a refund correctly once but mishandles it the next four times is not reliable. To measure true consistency, ThinkingBox runs each task 20 independent times from an identical clean backend environment. The benchmark reports three distinct metrics:

  • pass@1: The share of all attempts that succeeded, answering "How does it usually do?"
  • pass@20: The share of tasks solved at least once in 20 tries, answering "Can it ever do this?" (Breadth).
  • Observed 20/20: The literal count of tasks that passed all 20 recorded attempts, answering "Can it always be correct?"

Benchmark Results

ThinkingBox-Bench evaluates 507 stateful business workflows across five domains: Retail, Auto insurance, Travel, Neobank, and Consulting. Under the pass@1 metric, the results reveal a clear hierarchy:

  • Proprietary Models: Claude Opus 5.5 leads the overall task-weighted score at 67.16%, slightly ahead of Claude Opus 5 (66.50%) and GPT-5.4 (65.36%).
  • Open-Weight Models: Kimi-K3 is the strongest performer at 57.37%, closely trailing proprietary models like GPT-6 Astra.

The results demonstrate that domain significantly impacts performance. Repetition is the ultimate trust test for AI agents. Users can run the benchmark themselves through OpenEnv to evaluate their own models.

Source