Allen Institute for AI (AI2) has introduced BenchMIRT, a new methodology that audits large language model benchmarks at the granular level of individual questions and tasks. Using multidimensional Item Response Theory—a technique borrowed from psychometrics—BenchMIRT analyzes how 100 different LLMs performed across 16 benchmarks comprising more than 34,000 questions to identify what capabilities are actually being measured. Applied to six general reasoning benchmarks and ten safety-focused evaluations, BenchMIRT independently discovered that models' benchmark performance breaks down into two dominant capability dimensions: safety and general reasoning. However, the analysis revealed that many individual benchmarks measure more than their stated purpose. For example, BBQ, designed to test social bias, actually correlates more strongly with general reasoning ability than safety behavior. Similarly, WMDP, which evaluates dangerous dual-use knowledge, measures general reasoning more than safety—and interestingly, stronger reasoning associates with lower WMDP scores since refusal is the desired response. The research shows that single benchmark scores often combine multiple confounded signals, potentially misleading researchers about true model capabilities. HarmBench demonstrates that while harmful content and contextual questions align with safety, copyright-related questions measure general reasoning instead. BenchMIRT's methodology enables researchers to disentangle these mixed signals and gain clearer insight into what drives benchmark performance, offering a more nuanced approach to LLM evaluation.