All benchmarks
MathHard set

FrontierMath (Tiers 1–3)

290 unpublished research-level math problems written by professional mathematicians, tiers 1–3. Epoch AI runs it.

Competition and research mathematics, graded on the final answer or the proof. One of 5 comparable tests in math, a category the Index averages.

Read with care. Commissioned by OpenAI, which has access to much of the problem set; Epoch discloses this and keeps a holdout it does not share.

Publisher: Epoch AI (internal runs)

What a model is asked to do

Solve unpublished research-level mathematics problems written by professional mathematicians (tiers 1 to 3).

For exampleDetermine how many distinct values a given arithmetic function takes over the first million integers, with a verifiable argument.

Why it matters. Problems that take a working mathematician hours. A score here is a score against the frontier of the field.

The example is original and illustrative, not an item from the dataset.

Models scored here
61
Who produced the numbers
Run by Epoch AI
Items graded
n = 290 (0.3% each)
Best published result
93.7%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (93.7%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.

Snapshot September 8, 2026

Ranking on this test

61 models on FrontierMath (Tiers 1–3)

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.3 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-6 AstraOpenAI93.7%Epoch-run
2Claude Fable 5.1Anthropic90.2%Epoch-run
3GPT-5.6 SolOpenAI89.1%Epoch-run
4GPT-5.5 ProOpenAI87.7%Epoch-run
5Claude Fable 5Anthropic87.0%Epoch-run
6GPT-5.6 TerraOpenAI86.0%Epoch-run
7Claude Opus 5Anthropic85.6%Epoch-run
8GPT-5.5OpenAI85.3%Epoch-run
9GPT-5.4 ProOpenAI82.5%Epoch-run
10GPT-5.6 LunaOpenAI82.1%Epoch-run
11Claude Opus 4.8Anthropic80.0%Epoch-run
12GPT-5.4OpenAI78.6%Epoch-run
13Qwen3.8 MaxAlibaba74.7%Epoch-run
14GPT-5.2 ProOpenAI74.0%Epoch-run
15Kimi K3Moonshot AI72.2%Epoch-run
16Gemini 3.7 FlashGoogle71.6%Epoch-run
17Claude Opus 4.7Anthropic70.2%Epoch-run
18GLM 5.3Zhipu AI68.8%Epoch-run
19GPT-5.2OpenAI67.4%Epoch-run
20Claude Opus 4.6Anthropic66.0%Epoch-run
21Grok 4.6xAI66.0%Epoch-run
22Claude Sonnet 5Anthropic65.6%Epoch-run
23DeepSeek V4 ProDeepSeek64.6%Epoch-run
24Qwen3.7 MaxAlibaba64.6%Epoch-run
25Gemini 3.5 FlashGoogle62.8%Epoch-run
26Gemini 3.1 ProGoogle59.6%Epoch-run
27GLM 5.2Zhipu AI59.2%Epoch-run
28Gemini 3.6 FlashGoogle58.9%Epoch-run
29DeepSeek V4 FlashDeepSeek57.5%Epoch-run
30Grok 4.5xAI57.2%Epoch-run
31Kimi K2.6Moonshot AI57.2%Epoch-run
32GPT-5 ProOpenAI55.8%Epoch-run
33GLM 5.3 FlashZhipu AI55.8%Epoch-run
34GPT-5OpenAI55.4%Epoch-run
35Kimi K2.7 CodeMoonshot AI54.0%Epoch-run
36Gemini 3 FlashGoogle51.2%Epoch-run
37GPT-5.4 miniOpenAI51.2%Epoch-run
38GPT-5 miniOpenAI46.7%Epoch-run
39Inkling SmallThinking Machines46.3%Epoch-run
40Grok 4.20xAI44.9%Epoch-run
41GPT-5.4 nanoOpenAI44.9%Epoch-run
42Grok 4.3xAI42.8%Epoch-run
43Qwen3.6 PlusAlibaba38.2%Epoch-run
44GLM 5.1Zhipu AI36.8%Epoch-run
45Qwen3.6 27BAlibaba35.1%Epoch-run
46Claude Opus 4.5Anthropic34.4%Epoch-run
47Qwen3.7 PlusAlibaba34.4%Epoch-run
48InklingThinking Machines33.3%Epoch-run
49Qwen3.5 397B-A17BAlibaba31.2%Epoch-run
50Gemini 3.1 Flash-LiteGoogle27.7%Epoch-run
51GPT-5.5 InstantOpenAI26.3%Epoch-run
52Gemini 3.5 Flash-LiteGoogle26.0%Epoch-run
53Gemini 2.5 ProGoogle24.6%Epoch-run
54Claude Sonnet 4.5Anthropic23.9%Epoch-run
55Qwen3.6 FlashAlibaba22.5%Epoch-run
56Qwen3.6 35B-A3BAlibaba20.4%Epoch-run
57GPT-5 nanoOpenAI20.0%Epoch-run
58Qwen3.7 FlashAlibaba19.3%Epoch-run
59Qwen3 MaxAlibaba18.9%Epoch-run
60Qwen3.5 FlashAlibaba18.2%Epoch-run
61Claude Opus 4.1Anthropic12.6%Epoch-run
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.