GPQA Diamond
198 PhD-level four-option questions in physics, chemistry and biology, written to be Google-proof. Epoch AI runs most models itself 16 times; random guessing scores 25%.
Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.
Read with care. Frontier models cluster in the eighties and nineties; Epoch's standard errors on this test are 2 to 3 points.
Publisher: Epoch AI (internal runs)What a model is asked to do
Answer a PhD-level multiple-choice question in physics, chemistry or biology that a web search will not settle.
For exampleA sealed flask of an ideal gas is heated until its pressure doubles at constant volume. Which of the four statements about the distribution of molecular speeds is true?
Why it matters. Expert-level science questions are the closest thing to asking a specialist colleague. A model that gets them right can check a technical claim instead of echoing it.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 78
- Who produced the numbers
- Run by Epoch AI
- Items graded
- n = 198 (0.5% each)
- Best published result
- 95.8%
- Chance level
- 25%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is measured as a share of the best published result (95.8%) above chance: random guessing (25%) reads 0 and the frontier reads 100, so a high-guessing exam is not flattered.
- Provenance
- Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.
Snapshot September 8, 2026
Ranking on this test
78 models on GPQA Diamond
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, measured above chance). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.5 points, so read gaps smaller than that as noise.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.