A new benchmark co-developed by Microsoft and Hugging Face has exposed a critical measurement gap in how AI agents are evaluated. ThinkingBox measures whether agents actually change backend databases correctly, rather than just checking if they make valid tool calls or generate appropriate responses. The finding is stark: agents routinely appear successful while leaving databases in incorrect states. In one example, a customer service agent correctly documented a refund but failed to keep the underlying shipping exception open as required—yet traditional evaluation methods would mark it as successful. The benchmark runs 507 business-critical tasks 20 times each across multiple LLM models. Results reveal the scale of the reliability problem: of 121,680 valid trials, nearly 80,000 failed their actual objective, yet 67% of those failures still terminated cleanly with no reported errors. Among failures, 77.6% involved incorrect database values, 43% created unintended side effects, and 25% missed required changes. The data highlights why traditional success metrics are misleading—they measure appearance, not outcome. Claude Opus 5.5 currently leads overall at 67.16% pass@1, with open-weight models like Kimi-K3 showing competitive performance. However, the most revealing metric is observed 20/20—tasks that succeeded all 20 times—which shows even top performers struggle with consistency. The benchmark is now available through Hugging Face and OpenEnv, offering enterprises a way to evaluate AI agents on actual reliability rather than proxies for competence.