FrontierMath Erdős 2026: Epoch AI Lean Eval of 68 Open Problems
Epoch posted FrontierMath Erdős on 1 September: 68 open problems, Lean-checked. The story is the protocol, not another flagship.
On 1 September 2026 Epoch AI published FrontierMath Erdős: 68 still-open Erdős problems that a curator judged mathematically serious, stated in Lean. A system scores only if it returns a machine-checked proof or disproof inside a fixed budget.
This is not another flagship recap. The piece stays with Epoch’s own problem set, protocol and first table: a pre-release GPT-6 Astra scored 3% (2 of 68); four other systems scored 0% on the same budget.
What the benchmark actually measures
“An Erdős problem” is not one difficulty class. Thomas Bloom’s erdosproblems.com then listed about 1,217 problems, 652 still open. Epoch asked Bloom to pick the 68 he found both important and hard — roughly 10% of the unsolved set, all open as of August 2026.
Curation, not the whole catalogue
Bloom roughly estimated that only 3–5 Erdős problems of this caliber had been solved by AI as of August 2026. The list is subjective; Epoch writes that the job is done if some readers object to a few inclusions and still recognise others as obvious.
Lean checks, not a natural-language paper
A solve is a Lean proof or disproof that passes verification. Fifty of the 68 were already formalised in Google’s Formal Conjectures project; the remaining 18 were formalised with AI. Epoch says it is not a Lean expert shop and is seeking further review. Checking also uses Comparator from the Lean FRO, built to resist cheating submissions.
Open harness, fixed budget
The evaluation code is open-sourced. The default is one attempt per problem, a $300 inference budget and 72 hours of working time. Models have no internet; they get an offline collection of mathematics papers plus tools such as a computer algebra system.
How the first run was scored
-
1
Start with the fixed protocol
Five systems have been run: a pre-release GPT-6 Astra, GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5. One attempt each. A problem counts only when Lean verification passes.
-
2
Then the official scores
Only Astra solved anything: 2 of 68, or 3%. It disproved problem 74 with a counterexample ($218, 15 hours) and proved problem 126 ($247, 16 hours). Astra’s other 66 attempts, and every attempt by the other four systems, exhausted the budget without a verified proof.
-
3
Extra attempts are not this score
Separately, Epoch ran looser attempts on the same pre-release Astra with larger budgets and varied scaffolds. Those runs solved 5 problems at least once (also 1, 548 and 571), at more than $220,000 versus about $20,000 for the official run. Epoch is explicit: the FrontierMath Erdős score remains 3%.
What 3% does and does not say
Epoch’s own gloss: in this scaffold, at this budget, Astra solved about 3% of a list of significant open problems. That is a real data point on AI math breakthroughs — and also that an arbitrary problem of this class is still unlikely to fall, at least if the result must be formalised inside $300.
| System | Official score | How to read it |
|---|---|---|
| GPT-6 Astra | 3% (2/68) | Pre-release; only official solves |
| GPT-5.6 Sol / GPT-5.5 | 0% | No verified proof in budget |
| Claude Fable 5.1 / 5 | 0% | Same protocol, not a second set |
Epoch writes that the flashiest AI math results have come from inside labs, with little public method. FrontierMath Erdős is meant to name the problems, the harness and the budget in one place.
- Formalisation is a second project
- A natural-language breakthrough and a Lean formalisation are almost separate jobs. Epoch cites the unit-distance conjecture: an 18-page write-up versus a later Lean effort of about 1.2 million lines, mostly to rebuild a deep result not yet in the library.
- Contamination will grow
- For now the bench is run as a classical public set, without strong contamination guards. None of the 68 had a known solution as of August 2026, so a model whose training ends before then cannot have memorised one. Later models can be compared on the still-open remainder.
- Erdős is not all of mathematics
- The list is not a sample of every field. Epoch expects some correlation with general math progress and says the inference is not airtight.
When the team needs one table
After the post, groups usually want one picture: the official 3% and 0% columns, and why extra spend does not rewrite the score. No extra meeting suite. Open a short space on tidemeet or see how to create a space. For prompt-level audit, read BenchMIRT; for how a harness changes a published score, see ARC-AGI-3’s two scores.
# protocol
lean-proof | $300 | 72h
official = 2/68
extras ≠ score
Questions
Is this a GPT-6 Astra launch story?
No. The subject is Epoch’s new eval protocol. Astra is one of five systems in the first run, and the copy Epoch scored was a pre-release.
Why don’t the three extra solves count?
Those attempts used larger budgets, varied scaffolds and uneven retry counts. Epoch excludes them from the FrontierMath Erdős score.
Does 0% mean the other models cannot do math?
No. 0% means that under $300, 72 hours and one attempt, they did not return a Lean-verified proof. Most attempts ran out of budget rather than submitting a failed proof.
What does this have to do with a temporary meeting tool?
No product tie. If you only need one sketch of the official table versus the informal extra runs, temporary meeting tools says when a short browser space is enough.