ARC-AGI-3 two scores explained: Astra Standard vs Adapter 2026

ARC Prize posted 62.7% and 99.9% for Astra on ARC-AGI-3 on 3 September. The story is two harnesses, not another flagship recap.

On 3 September 2026 the ARC Prize Foundation posted two ARC-AGI-3 Semi-Private scores for GPT-6 Astra: 62.7% on the Standard harness and 99.9% on the Provider Adapter. Same model, same bench, more than thirty points apart.

This is not another flagship launch recap. OpenAI was rolling Astra out the same day. This piece reads only the two harnesses ARC Prize published, and the foundation’s own rule: Standard for comparing models, Adapter for a provider’s own context management.

What was measured Two harnesses How to read the scores

Two harnesses, two questions

ARC-AGI-3 is an interactive reasoning benchmark: agents must explore unfamiliar turn-based environments, infer goals, and build an internal model with no instruction sheet. The foundation has said humans can solve every environment; when the bench launched, every frontier system it tested scored below 1%.

01

Standard harness

A minimal, provider-neutral interface. The model may carry forward only the visible notes it chooses to keep. Astra (max) scored 62.7% on Semi-Private for $26,098. The foundation says a future AGI should clear the bench under these conditions, and uses this setup for cross-provider comparison.

02

Provider Adapter

Preserves opaque reasoning state between requests and uses compaction on long conversations so the model can reuse prior work. Astra (high) scored 99.9% for $18,817. Higher score, lower bill.

03

Both go on the board

ARC Prize says it will report both Standard and Provider Adapter results on the ARC-AGI leaderboard, each condition labeled. The open test repo and the testing policy cover both.

How to read the two numbers

  1. 1

    Start with Standard

    62.7% is the figure you can set beside other models. R&D World and later explainers treat it as the buyer’s number. OpenAI’s launch page leads with 99.9%.

  2. 2

    Then ask what the Adapter measures

    The Adapter asks how far the model goes with the context-management features its provider designed for it. Across Public and Semi-Private and every reasoning level, Adapter runs were about 3.66× faster in recorded elapsed time and used 49% fewer tokens on the 167 game–reasoning pairs both harnesses solved.

  3. 3

    Then look at action efficiency

    On the Provider Adapter, Astra (max) used fewer actions than the median human baseline on 96.0% of levels and 51.7% fewer actions per level on average. The foundation calls that a human-parity milestone on action efficiency, not a completion-only score.

62.7%
Standard Semi-Private (max)
99.9%
Provider Adapter (high)
96%
Levels under human-median actions

Where the foundation draws the line

In Astra’s replays the model compresses unfamiliar mechanics into a compact symbolic world model and invents shorthand to track state and plans. In the PRO-LONG red-team harness it also wrote per-game parsers and planners; the foundation warns that is model-plus-tools, not comparable to the controlled human test.

Item Public figure How to read it
Standard / max 62.7%, $26,098 Use this to compare providers
Adapter / high 99.9%, $18,817 With provider context management
Adapter / max 98.6%, $17,332 Higher effort is not always a higher score

ARC Prize’s own line: saturating ARC-AGI-3 is not “proof of achieving AGI.” The environments are bounded, the mechanics deterministic, the goals closed-ended — not the open world.

What Standard keeps
The model decides what to store in visible notes. The interface is minimal so scores can sit side by side.
What the Adapter keeps
Opaque reasoning state survives between requests; long chats are compacted. Score and wall-clock both move.
How the human baseline was built
About 500 members of the public before launch; the per-level baseline is the median action count among people who finished. Participants were paid per session, not recruited as puzzle specialists.

When the team needs one sketch

After the post, groups usually want one picture: Standard 62.7%, Adapter 99.9%, and the “not AGI” line. No extra meeting suite. Open a short space on tidemeet or see how to create a space. For another independent eval the same week, read METR’s OpenAI / Hugging Face report; for a lab protocol, see DeepMind’s double-blind enclave.

# axis
standard-62.7 → adapter-99.9 → not-agi

Questions

Is this a GPT-6 Astra launch recap?

No. Astra began a staged rollout on 3 September. This piece only explains the two ARC-AGI-3 scores ARC Prize posted that day and why the harnesses differ.

Which score should I use against other models?

The foundation treats Standard as the apples-to-apples condition across providers. The Adapter score is an upper bound with that provider’s context management; do not press it against models that did not get the same adapter.

Does 99.9% mean AGI?

The foundation says no. ARC-AGI-3 environments are bounded and closed-ended. They call Astra meaningful progress toward generalization and stop there.

What does this have to do with a temporary meeting tool?

No product tie. If you only need one sketch of the two harnesses, temporary meeting tools says when a short browser space is enough.

Create a free space