All benchmarks
ReasoningInternally runnable

GPQA Diamond

198 PhD-level four-option questions in physics, chemistry and biology, written to be Google-proof. Epoch AI runs most models itself 16 times; random guessing scores 25%.

Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.

Read with care. Frontier models cluster in the eighties and nineties; Epoch's standard errors on this test are 2 to 3 points.

Publisher: Epoch AI (internal runs)

What a model is asked to do

Answer a PhD-level multiple-choice question in physics, chemistry or biology that a web search will not settle.

For exampleA sealed flask of an ideal gas is heated until its pressure doubles at constant volume. Which of the four statements about the distribution of molecular speeds is true?

Why it matters. Expert-level science questions are the closest thing to asking a specialist colleague. A model that gets them right can check a technical claim instead of echoing it.

The example is original and illustrative, not an item from the dataset.

Models scored here
78
Who produced the numbers
Run by Epoch AI
Items graded
n = 198 (0.5% each)
Best published result
95.8%
Chance level
25%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is measured as a share of the best published result (95.8%) above chance: random guessing (25%) reads 0 and the frontier reads 100, so a high-guessing exam is not flattered.
Provenance
Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.

Snapshot September 8, 2026

Ranking on this test

78 models on GPQA Diamond

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, measured above chance). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.5 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-6 AstraOpenAI95.8%Epoch-run
2Gemini 3.7 FlashGoogle94.8%Epoch-run
3GPT-5.4 ProOpenAI94.6%Epoch-run
4Gemini 3.1 ProGoogle94.4%Epoch-run
5Gemini 3.6 FlashGoogle94.1%Epoch-run
6Grok 4.6xAI94.0%Epoch-run
7Claude Opus 5Anthropic93.9%Epoch-run
8GPT-5.6 SolOpenAI93.5%Epoch-run
9Grok 4.5xAI93.4%Epoch-run
10GPT-5.6 TerraOpenAI93.3%Epoch-run
11GPT-5.4OpenAI93.3%Epoch-run
12Kimi K3Moonshot AI93.1%Epoch-run
13Gemini 3.5 FlashGoogle92.8%Epoch-run
14Qwen3.8 MaxAlibaba92.7%Epoch-run
15Gemini 3 ProGoogle92.6%Epoch-run
16GLM 5.2Zhipu AI91.9%Epoch-run
17DeepSeek V4 ProDeepSeek91.7%Epoch-run
18GPT-5.6 LunaOpenAI91.6%Epoch-run
19GPT-5.2OpenAI91.4%Epoch-run
20Claude Opus 4.8Anthropic91.0%Epoch-run
21DeepSeek V4 FlashDeepSeek91.0%Epoch-run
22Qwen3.7 MaxAlibaba90.9%Epoch-run
23GLM 5.3Zhipu AI90.9%Epoch-run
24MiniMax M3MiniMax90.9%Epoch-run
25Kimi K2.6Moonshot AI90.8%Epoch-run
26GPT-5.5OpenAI90.7%Epoch-run
27Claude Opus 4.6Anthropic90.5%Epoch-run
28Claude Sonnet 5Anthropic90.5%Epoch-run
29Claude Opus 4.7Anthropic90.2%Epoch-run
30GLM 5.3 FlashZhipu AI90.2%Epoch-run
31GLM 5.1Zhipu AI89.9%Epoch-run
32Muse SparkMeta89.8%Epoch-run
33Gemini 3 FlashGoogle89.4%Epoch-run
34Grok 4.20xAI89.3%Epoch-run
35Grok 4.3xAI88.8%Epoch-run
36Inkling SmallThinking Machines88.5%Epoch-run
37Qwen3.6 PlusAlibaba88.4%Epoch-run
38InklingThinking Machines88.3%Epoch-run
39Kimi K2.7 CodeMoonshot AI87.9%Epoch-run
40Qwen3.7 PlusAlibaba87.9%Epoch-run
41GLM 5Zhipu AI87.8%Epoch-run
42GPT-5.1OpenAI87.6%Epoch-run
43Kimi K2.5Moonshot AI87.6%Epoch-run
44Claude Sonnet 4.6Anthropic87.4%Epoch-run
45Qwen3.6 MaxAlibaba87.4%Epoch-run
46Grok 4xAI87.0%Epoch-run
47GPT-5.4 miniOpenAI86.9%Epoch-run
48Qwen3.5 397B-A17BAlibaba86.4%Epoch-run
49GPT-5OpenAI86.2%Epoch-run
50Claude Opus 4.5Anthropic86.0%Epoch-run
51Claude Fable 5Anthropic85.9%Epoch-run
52Qwen3.6 27BAlibaba85.9%Epoch-run
53Nemotron 3 UltraNVIDIA85.4%Epoch-run
54Gemini 2.5 ProGoogle85.3%Epoch-run
55Qwen3.5 PlusAlibaba84.8%Epoch-run
56Qwen3.6 35B-A3BAlibaba84.8%Epoch-run
57Kimi K2Moonshot AI84.2%Epoch-run
58Qwen3.5 35B-A3BAlibaba83.5%Epoch-run
59DeepSeek V3.2DeepSeek83.4%Epoch-run
60GLM 4.7Zhipu AI83.3%Epoch-run
61Qwen3.6 FlashAlibaba83.3%Epoch-run
62Gemini 3.5 Flash-LiteGoogle83.3%Epoch-run
63GPT-5.5 InstantOpenAI82.5%Epoch-run
64Claude Sonnet 4.5Anthropic82.3%Epoch-run
65Qwen3.7 FlashAlibaba82.3%Epoch-run
66Qwen3.5 FlashAlibaba82.3%Epoch-run
67Gemini 3.1 Flash-LiteGoogle81.8%Epoch-run
68Qwen3.5 9BAlibaba79.0%Epoch-run
69GPT-5.4 nanoOpenAI78.5%Epoch-run
70Claude Opus 4.1Anthropic77.3%Epoch-run
71Gemma 4 31BGoogle75.8%Epoch-run
72GPT-OSS 120BOpenAI75.8%Epoch-run
73GPT-5 miniOpenAI75.0%Epoch-run
74Gemma 4 26B A4BGoogle73.2%Epoch-run
75Qwen3 MaxAlibaba72.6%Epoch-run
76Claude Haiku 4.5Anthropic71.2%Epoch-run
77GPT-5 nanoOpenAI69.4%Epoch-run
78GPT-OSS 20BOpenAI60.8%Epoch-run
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.