All benchmarks
Reasoning

CritPt

Unpublished research-level physics problems graded by an official server — accuracy on the full set, as run by Artificial Analysis.

Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.

Publisher: CritPt / Artificial Analysis

What a model is asked to do

Solve an unpublished research-level physics problem to a numeric or symbolic answer checked by an official server.

For exampleDerive the leading correction to the decay rate of a metastable state coupled to a bath with the given spectral density, and give it in closed form.

Why it matters. Research physics is far past textbook recall. Only a few models produce anything a physicist would accept.

The example is original and illustrative, not an item from the dataset.

Models scored here
78
Who produced the numbers
Official leaderboard
Best published result
32.3%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (32.3%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

78 models on CritPt

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1GPT-5.6 SolOpenAI32.3%Leaderboard
2GPT-6 AstraOpenAI31.7%Leaderboard
3Claude Fable 5.1Anthropic31.1%Leaderboard
4GPT-5.5 ProOpenAI30.6%Leaderboard
5GPT-5.6 TerraOpenAI30.0%Leaderboard
6GPT-5.4 ProOpenAI30.0%Leaderboard
7Claude Opus 5Anthropic29.1%Leaderboard
8Claude Fable 5Anthropic28.6%Leaderboard
9GPT-5.5OpenAI27.1%Leaderboard
10Gemini 3 Deep ThinkGoogle25.7%Leaderboard
11GPT-5.4OpenAI23.4%Leaderboard
12Kimi K3Moonshot AI23.4%Leaderboard
13Claude Opus 4.8Anthropic20.9%Leaderboard
14GLM 5.2Zhipu AI20.9%Leaderboard
15GPT-5.6 LunaOpenAI20.6%Leaderboard
16Qwen3.8 MaxAlibaba20.0%Leaderboard
17Grok 4.6xAI19.7%Leaderboard
18GLM 5.3Zhipu AI19.1%Leaderboard
19Gemini 3.8 FlashGoogle18.3%Leaderboard
20DeepSeek V4 ProDeepSeek18.0%Leaderboard
21Gemini 3.1 ProGoogle17.7%Leaderboard
22Muse Spark 1.2Meta17.7%Leaderboard
23Claude Sonnet 5Anthropic16.9%Leaderboard
24DeepSeek V4 FlashDeepSeek16.6%Leaderboard
25Grok 4.5xAI15.4%Leaderboard
26GLM 5.3 FlashZhipu AI15.4%Leaderboard
27Muse Spark 1.1Meta15.1%Leaderboard
28Gemini 3.7 FlashGoogle14.3%Leaderboard
29Qwen3.7 MaxAlibaba13.4%Leaderboard
30Gemini 3.5 FlashGoogle13.1%Leaderboard
31GPT-5OpenAI12.6%Leaderboard
32Claude Opus 4.7Anthropic12.0%Leaderboard
33Muse SparkMeta11.3%Leaderboard
34Gemini 3.6 FlashGoogle10.6%Leaderboard
35Kimi K2.7 CodeMoonshot AI10.0%Leaderboard
36GPT-5.4 miniOpenAI10.0%Leaderboard
37GPT-5.4 nanoOpenAI9.3%Leaderboard
38Qwen3.7 PlusAlibaba9.1%Leaderboard
39Inkling SmallThinking Machines8.3%Leaderboard
40Kimi K2.6Moonshot AI8.0%Leaderboard
41Grok 4.3xAI8.0%Leaderboard
42Gemini 3 ProGoogle6.9%Leaderboard
43InklingThinking Machines5.4%Leaderboard
44Qwen3.8 27BAlibaba5.4%Leaderboard
45GPT-5.1OpenAI4.9%Leaderboard
46GLM 5.1Zhipu AI4.6%Leaderboard
47MiMo V2.5 ProXiaomi4.0%Leaderboard
48MiniMax M3MiniMax3.7%Leaderboard
49MiMo V2.5Xiaomi3.7%Leaderboard
50Claude Sonnet 4.6Anthropic3.1%Leaderboard
51Kimi K2.5Moonshot AI3.1%Leaderboard
52Nemotron 3 UltraNVIDIA3.1%Leaderboard
53Nemotron 3 SuperNVIDIA3.1%Leaderboard
54Qwen3.6 PlusAlibaba2.9%Leaderboard
55DeepSeek V3.2DeepSeek2.9%Leaderboard
56Gemini 2.5 ProGoogle2.0%Leaderboard
57GLM 4.7Zhipu AI1.7%Leaderboard
58DeepSeek V3.1 TerminusDeepSeek1.7%Leaderboard
59Gemma 4 31BGoogle1.4%Leaderboard
60GPT-OSS 20BOpenAI1.4%Leaderboard
61Claude Sonnet 4.5Anthropic1.1%Leaderboard
62GLM 4.6Zhipu AI1.1%Leaderboard
63Gemini 3.1 Flash-LiteGoogle1.1%Leaderboard
64GPT-OSS 120BOpenAI1.1%Leaderboard
65Gemini 2.5 FlashGoogle1.1%Leaderboard
66Qwen3.6 27BAlibaba0.9%Leaderboard
67Qwen3.5 122B-A10BAlibaba0.9%Leaderboard
68Qwen3.5 35B-A3BAlibaba0.6%Leaderboard
69MiniMax M2.7MiniMax0.6%Leaderboard
70Qwen3.6 35B-A3BAlibaba0.3%Leaderboard
71Qwen3.5 9BAlibaba0.3%Leaderboard
72Gemma 4 26B A4BGoogle0.0%Leaderboard
73Mistral Large 3Mistral AI0.0%Leaderboard
74GPT-5 miniOpenAI0.0%Leaderboard
75Mistral Medium 3.5Mistral AI0.0%Leaderboard
76GPT-5.5 InstantOpenAI0.0%Leaderboard
77Claude Haiku 4.5Anthropic0.0%Leaderboard
78Gemini 3.5 Flash-LiteGoogle0.0%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.