ProofBench
Full written proofs, not final answers, graded for rigor.
Competition and research mathematics, graded on the final answer or the proof. One of 5 comparable tests in math, a category the Index averages.
Publisher: Vals AIWhat a model is asked to do
Write a complete, rigorous proof, graded for rigor rather than for a final answer.
For exampleProve that every sequence of real numbers has a monotone subsequence, and say exactly where completeness is or is not used.
Why it matters. A proof shows the reasoning itself. A right answer with a wrong argument scores nothing.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 61
- Who produced the numbers
- Official leaderboard
- Best published result
- 100.0%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (100.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
61 models on ProofBench
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | Claude Fable 5.1Anthropic | 100.0% | 100.0 | max | Leaderboard |
| 2 | GPT-6 AstraOpenAI | 99.0% | 99.0 | — | Leaderboard |
| 3 | Claude Opus 5Anthropic | 99.0% | 99.0 | max | Leaderboard |
| 4 | Claude Fable 5Anthropic | 95.0% | 95.0 | max · reasoning effort: max | Leaderboard |
| 5 | Kimi K3Moonshot AI | 87.0% | 87.0 | — | Leaderboard |
| 6 | GPT-5.6 SolOpenAI | 83.0% | 83.0 | max · reasoning effort: max | Leaderboard |
| 7 | Claude Sonnet 5Anthropic | 77.0% | 77.0 | max · reasoning effort: max | Leaderboard |
| 8 | GPT-5.6 TerraOpenAI | 74.0% | 74.0 | xhigh · reasoning effort: xhigh | Leaderboard |
| 9 | Claude Opus 4.8Anthropic | 69.0% | 69.0 | max · reasoning effort: max | Leaderboard |
| 10 | GPT-5.6 LunaOpenAI | 60.0% | 60.0 | max · reasoning effort: max | Leaderboard |
| 11 | Gemini 3.7 FlashGoogle | 58.0% | 58.0 | — | Leaderboard |
| 12 | Qwen3.8 MaxAlibaba | 58.0% | 58.0 | max | Leaderboard |
| 13 | GPT-5.4OpenAI | 56.0% | 56.0 | 2026-03-05 · xhigh · reasoning effort: xhigh | Leaderboard |
| 14 | DeepSeek V4 FlashDeepSeek | 56.0% | 56.0 | — | Leaderboard |
| 15 | Claude Opus 4.7Anthropic | 54.0% | 54.0 | max · reasoning effort: max | Leaderboard |
| 16 | Grok 4.6xAI | 51.0% | 51.0 | — | Leaderboard |
| 17 | GPT-5.5OpenAI | 50.0% | 50.0 | xhigh · reasoning effort: xhigh | Leaderboard |
| 18 | Claude Opus 4.6Anthropic | 50.0% | 50.0 | max · reasoning effort: max | Leaderboard |
| 19 | DeepSeek V4 ProDeepSeek | 50.0% | 50.0 | — | Leaderboard |
| 20 | GLM 5.3Zhipu AI | 49.0% | 49.0 | max | Leaderboard |
| 21 | Gemini 3.8 FlashGoogle | 48.0% | 48.0 | — | Leaderboard |
| 22 | Claude Sonnet 4.6Anthropic | 45.0% | 45.0 | max · reasoning effort: max | Leaderboard |
| 23 | Muse Spark 1.2Meta | 43.0% | 43.0 | — | Leaderboard |
| 24 | Muse Spark 1.1Meta | 39.0% | 39.0 | reasoning effort: xhigh | Leaderboard |
| 25 | Gemini 3.6 FlashGoogle | 36.0% | 36.0 | reasoning effort: high | Leaderboard |
| 26 | Claude Opus 4.5Anthropic | 36.0% | 36.0 | 20251101 · reasoning effort: high | Leaderboard |
| 27 | GLM 5.2Zhipu AI | 35.0% | 35.0 | max · reasoning effort: max | Leaderboard |
| 28 | Gemini 3.5 FlashGoogle | 31.0% | 31.0 | high · reasoning effort: high | Leaderboard |
| 29 | Grok 4.5xAI | 31.0% | 31.0 | high · reasoning effort: high | Leaderboard |
| 30 | Gemini 3.1 ProGoogle | 26.0% | 26.0 | preview · reasoning effort: high | Leaderboard |
| 31 | Qwen3.7 MaxAlibaba | 26.0% | 26.0 | max | Leaderboard |
| 32 | GLM 5.1Zhipu AI | 22.2% | 22.2 | — | Leaderboard |
| 33 | MiMo V2.5 ProXiaomi | 22.0% | 22.0 | — | Leaderboard |
| 34 | GPT-5.4 miniOpenAI | 21.0% | 21.0 | 2026-03-17 · xhigh · reasoning effort: xhigh | Leaderboard |
| 35 | GLM 5.3 FlashZhipu AI | 21.0% | 21.0 | max | Leaderboard |
| 36 | Gemini 3 ProGoogle | 20.0% | 20.0 | preview · reasoning effort: high | Leaderboard |
| 37 | Claude Sonnet 4.5Anthropic | 19.0% | 19.0 | 20250929 | Leaderboard |
| 38 | GPT-5OpenAI | 18.0% | 18.0 | 2025-08-07 · high · reasoning effort: high | Leaderboard |
| 39 | MiniMax M3MiniMax | 18.0% | 18.0 | — | Leaderboard |
| 40 | Muse SparkMeta | 17.0% | 17.0 | — | Leaderboard |
| 41 | Kimi K2.6Moonshot AI | 16.0% | 16.0 | — | Leaderboard |
| 42 | Qwen3.8 27BAlibaba | 16.0% | 16.0 | — | Leaderboard |
| 43 | MiMo V2.5Xiaomi | 16.0% | 16.0 | — | Leaderboard |
| 44 | GPT-5.2OpenAI | 15.0% | 15.0 | 2025-12-11 · xhigh · reasoning effort: xhigh | Leaderboard |
| 45 | Gemini 3 FlashGoogle | 15.0% | 15.0 | preview · reasoning effort: high | Leaderboard |
| 46 | Grok 4.20xAI | 14.0% | 14.0 | reasoning | Leaderboard |
| 47 | Gemini 3.5 Flash-LiteGoogle | 13.0% | 13.0 | reasoning effort: high | Leaderboard |
| 48 | GPT-5 nanoOpenAI | 12.0% | 12.0 | 2025-08-07 · high · reasoning effort: high | Leaderboard |
| 49 | Grok 4.3xAI | 11.0% | 11.0 | high · reasoning effort: high | Leaderboard |
| 50 | GPT-5 miniOpenAI | 9.0% | 9.0 | 2025-08-07 · high · reasoning effort: high | Leaderboard |
| 51 | Mistral Medium 3.5Mistral AI | 9.0% | 9.0 | reasoning effort: high | Leaderboard |
| 52 | GPT-5.1 CodexOpenAI | 9.0% | 9.0 | max · reasoning effort: high | Leaderboard |
| 53 | DeepSeek V3.2DeepSeek | 8.0% | 8.0 | reasoning effort: none | Leaderboard |
| 54 | Inkling SmallThinking Machines | 6.0% | 6.0 | reasoning effort: 0.99 | Leaderboard |
| 55 | GLM 4.7Zhipu AI | 6.0% | 6.0 | — | Leaderboard |
| 56 | GPT-5.4 nanoOpenAI | 5.0% | 5.0 | 2026-03-17 · high · reasoning effort: high | Leaderboard |
| 57 | Grok 4.1 FastxAI | 4.0% | 4.0 | reasoning | Leaderboard |
| 58 | MiniMax M2.5MiniMax | 4.0% | 4.0 | — | Leaderboard |
| 59 | MiniMax M2.7MiniMax | 3.0% | 3.0 | — | Leaderboard |
| 60 | Nemotron 3 UltraNVIDIA | 2.0% | 2.0 | — | Leaderboard |
| 61 | InklingThinking Machines | 0.0% | 0.0 | — | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.