BenchMIRT explained: Ai2 prompt-level benchmark audit 2026
Ai2 posted BenchMIRT on 1 September: audit benches at the prompt, not the headline score. The story is two axes, not another flagship.
On 1 September 2026 the Allen Institute for AI (Ai2) published BenchMIRT on Hugging Face: an audit that reads a benchmark at the level of each prompt, not only the headline score.
This is not another flagship recap. The piece stays with what Ai2 wrote down: 100 open-weight LLMs, 16 benchmarks, more than 34K questions, and two dimensions the method recovered without being told the labels — safety and general reasoning.
What the method actually does
A benchmark is usually sold as one ability: safety, reasoning, or instruction following. A single item can still depend on more than that. One BBQ question about a grandson and grandfather booking an Uber probes age stereotype — and also who-is-who tracking and evidence rather than assumption.
From IRT to MIRT
Item Response Theory comes from psychometrics: not every question tells you the same amount. Ai2 had already used single-dimension IRT in Fluid Benchmarking. BenchMIRT extends that to multidimensional IRT so several abilities on the same item can be separated.
Estimates on both sides
For a model it estimates strength on the capabilities reflected in the selected benches. For each question it estimates difficulty and how well the item separates stronger from weaker models on those capabilities.
No capability labels
Training used 100 LLMs, 16 benchmarks and more than 34K items. Six reasoning benches include MMLU-Pro, GPQA, MATH and BBH. Ten come from the Olmo 3 safety suite, including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP and XSTest. The team did not tell the method which bench measures what.
Which two axes came back
-
1
Start with the stable recovery
The method independently recovered two dominant dimensions: safety and general reasoning. Reruns from scratch found the same pair. High scores on most reasoning benches tracked reasoning; jailbreak and harmful-content benches tracked safety.
-
2
Then the mis-grouped items
BBQ is usually filed under safety, but aligned much more with general reasoning: a low score may partly mean the item was hard to parse, not only stereotype use. WMDP tests dangerous dual-use knowledge in biology, chemistry and cybersecurity and also tracked reasoning more than safety. Stronger general reasoning was associated with lower WMDP scores, because refusing or failing to supply the dangerous knowledge is the desired response.
-
3
Then splits inside one bench
WildJailbreak averages harmful jailbreaks with benign prompts that test over-refusal. On HarmBench, standard and contextual harmful requests aligned more with safety; copyright items — for example generating the lyrics of “What a Wonderful World” — aligned more with general reasoning.
Does a thinner set still hold the picture
After ranking items by how informative they were, keeping 10% of questions across those 16 benches generally preserved nearly the same stronger/weaker picture on the underlying safety or reasoning capability; keeping 50% often matched the full set even more closely. On held-out questions the method had not seen answered, it predicted correct/incorrect 79% of the time; a simpler “use the model’s overall bench average” guess was right 70% of the time.
| Finding | Public claim | How to read it |
|---|---|---|
| Dominant axes | Safety / general reasoning | Recovered without labels |
| BBQ, WMDP | Closer to reasoning | A safety-suite score is not safety behavior |
| Item trim | 10% mostly keeps the picture | Not every item carries equal information |
Ai2 is explicit: this does not necessarily mean the benches are flawed. It means one score can stack several signals, and BenchMIRT is built to pull them apart.
- Cost if you only want a rank
- If the goal is ranking models by predicted performance on randomly held-out items, the benchmark’s average score does slightly better than BenchMIRT. The method’s edge is the question-level picture, not the overall table.
- Risk of deleting the useful items
- The same estimates that flag the most informative safety questions could be used to drop them, so an unsafe model passes a weaker eval. The authors think the transparency is worth that risk and still name it.
- Axes follow the set
- A different mix of evaluations could surface different capabilities. The two-axis story is tied to these 16 benches.
When the team needs one sketch
After the post, groups usually want one picture: a safety axis, a reasoning axis, and why BBQ / WMDP are not pure safety scores. No extra meeting suite. Open a short space on tidemeet or see how to create a space. For how a harness changes a published score, read ARC-AGI-3’s two scores; for eval isolation failure, see METR and Hugging Face.
# axis
safety ↔ general-reasoning
bbq/wmdp ≠ safety-only
Questions
Is this a new flagship model story?
No. BenchMIRT is an evaluation-audit method, not a model release. The sample stops at open-weight models from March 2025 and earlier.
Can I score 2026 Astra or Fable with it?
The post says the models used to train and evaluate BenchMIRT were all released by March 2025, so the analysis does not cover newer generations. Do not paste the two axes onto flagships that were not in the sample.
Does 10% of the items replace the full bench?
The authors say that on these 16 benches, the most informative 10% generally kept the stronger/weaker picture. If you want predicted rank on randomly held-out items, the full-bench average still does slightly better.
What does this have to do with a temporary meeting tool?
No product tie. If you only need one sketch of the two axes and the mis-grouped benches, temporary meeting tools says when a short browser space is enough.