All benchmarks
MathInternally runnable

MATH Level 5

The hardest tier of the MATH competition dataset. Saturated at the frontier; useful for the long tail.

Competition and research mathematics, graded on the final answer or the proof. One of 5 comparable tests in math, a category the Index averages.

Read with care. Epoch flags likely training contamination on this set; treat small gaps as meaningless.

Publisher: Epoch AI (internal runs)

What a model is asked to do

Solve the hardest tier of the MATH competition dataset to an exact final answer.

For exampleFind all real x such that the distance from x to 3 plus the distance from x to 11 equals 14.

Why it matters. Multi-step algebra and geometry, graded only on the final answer.

The example is original and illustrative, not an item from the dataset.

Models scored here
6
Who produced the numbers
Run by Epoch AI
Best published result
98.1%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (98.1%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.

Snapshot September 8, 2026

Ranking on this test

6 models on MATH Level 5

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1GPT-5OpenAI98.1%Epoch-run
2GPT-5 miniOpenAI97.8%Epoch-run
3Claude Sonnet 4.5Anthropic97.7%Epoch-run
4Qwen3 MaxAlibaba97.1%Epoch-run
5Claude Haiku 4.5Anthropic96.4%Epoch-run
6GPT-5 nanoOpenAI95.2%Epoch-run
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.