BenchMIRT explained: Ai2 prompt-level benchmark audit 2026

Ai2 posted BenchMIRT on 1 September: audit benches at the prompt, not the headline score. The story is two axes, not another flagship.

On 1 September 2026 the Allen Institute for AI (Ai2) published BenchMIRT on Hugging Face: an audit that reads a benchmark at the level of each prompt, not only the headline score.

This is not another flagship recap. The piece stays with what Ai2 wrote down: 100 open-weight LLMs, 16 benchmarks, more than 34K questions, and two dimensions the method recovered without being told the labels — safety and general reasoning.

Who published What was measured How to read it

What the method actually does

A benchmark is usually sold as one ability: safety, reasoning, or instruction following. A single item can still depend on more than that. One BBQ question about a grandson and grandfather booking an Uber probes age stereotype — and also who-is-who tracking and evidence rather than assumption.

01

From IRT to MIRT

Item Response Theory comes from psychometrics: not every question tells you the same amount. Ai2 had already used single-dimension IRT in Fluid Benchmarking. BenchMIRT extends that to multidimensional IRT so several abilities on the same item can be separated.

02

Estimates on both sides

For a model it estimates strength on the capabilities reflected in the selected benches. For each question it estimates difficulty and how well the item separates stronger from weaker models on those capabilities.

03

No capability labels

Training used 100 LLMs, 16 benchmarks and more than 34K items. Six reasoning benches include MMLU-Pro, GPQA, MATH and BBH. Ten come from the Olmo 3 safety suite, including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP and XSTest. The team did not tell the method which bench measures what.

Which two axes came back

  1. 1

    Start with the stable recovery

    The method independently recovered two dominant dimensions: safety and general reasoning. Reruns from scratch found the same pair. High scores on most reasoning benches tracked reasoning; jailbreak and harmful-content benches tracked safety.

  2. 2

    Then the mis-grouped items

    BBQ is usually filed under safety, but aligned much more with general reasoning: a low score may partly mean the item was hard to parse, not only stereotype use. WMDP tests dangerous dual-use knowledge in biology, chemistry and cybersecurity and also tracked reasoning more than safety. Stronger general reasoning was associated with lower WMDP scores, because refusing or failing to supply the dangerous knowledge is the desired response.

  3. 3

    Then splits inside one bench

    WildJailbreak averages harmful jailbreaks with benign prompts that test over-refusal. On HarmBench, standard and contextual harmful requests aligned more with safety; copyright items — for example generating the lyrics of “What a Wonderful World” — aligned more with general reasoning.

100
Open-weight models
16
Benchmarks
34K+
Questions / tasks

Does a thinner set still hold the picture

After ranking items by how informative they were, keeping 10% of questions across those 16 benches generally preserved nearly the same stronger/weaker picture on the underlying safety or reasoning capability; keeping 50% often matched the full set even more closely. On held-out questions the method had not seen answered, it predicted correct/incorrect 79% of the time; a simpler “use the model’s overall bench average” guess was right 70% of the time.

Finding Public claim How to read it
Dominant axes Safety / general reasoning Recovered without labels
BBQ, WMDP Closer to reasoning A safety-suite score is not safety behavior
Item trim 10% mostly keeps the picture Not every item carries equal information

Ai2 is explicit: this does not necessarily mean the benches are flawed. It means one score can stack several signals, and BenchMIRT is built to pull them apart.

Cost if you only want a rank
If the goal is ranking models by predicted performance on randomly held-out items, the benchmark’s average score does slightly better than BenchMIRT. The method’s edge is the question-level picture, not the overall table.
Risk of deleting the useful items
The same estimates that flag the most informative safety questions could be used to drop them, so an unsafe model passes a weaker eval. The authors think the transparency is worth that risk and still name it.
Axes follow the set
A different mix of evaluations could surface different capabilities. The two-axis story is tied to these 16 benches.

When the team needs one sketch

After the post, groups usually want one picture: a safety axis, a reasoning axis, and why BBQ / WMDP are not pure safety scores. No extra meeting suite. Open a short space on tidemeet or see how to create a space. For how a harness changes a published score, read ARC-AGI-3’s two scores; for eval isolation failure, see METR and Hugging Face.

# axis
safety ↔ general-reasoning
bbq/wmdp ≠ safety-only

Questions

Is this a new flagship model story?

No. BenchMIRT is an evaluation-audit method, not a model release. The sample stops at open-weight models from March 2025 and earlier.

Can I score 2026 Astra or Fable with it?

The post says the models used to train and evaluate BenchMIRT were all released by March 2025, so the analysis does not cover newer generations. Do not paste the two axes onto flagships that were not in the sample.

Does 10% of the items replace the full bench?

The authors say that on these 16 benches, the most informative 10% generally kept the stronger/weaker picture. If you want predicted rank on randomly held-out items, the full-bench average still does slightly better.

What does this have to do with a temporary meeting tool?

No product tie. If you only need one sketch of the two axes and the mis-grouped benches, temporary meeting tools says when a short browser space is enough.

Create a free space