All benchmarks
KnowledgeBeside the IndexInternally runnable

SimpleQA Verified

A thousand verified short-answer factual questions — how often the model is right, without being allowed to search. Epoch AI runs it.

Recall of facts without hallucinating — short-answer accuracy. Knowledge is a single-test category in this snapshot (fewer than 2 comparable tests), so it is shown beside the Index and never averaged into it.

Publisher: Epoch AI (internal runs)

What a model is asked to do

Answer a short factual question exactly, without searching, or say you do not know.

For exampleWhich architect designed the smaller second station that replaced the city's original 1901 harbor terminus?

Why it matters. How often a model is right about facts, and how rarely it invents one, decides whether you can trust it unsupervised.

The example is original and illustrative, not an item from the dataset.

Models scored here
57
Who produced the numbers
Run by Epoch AI
Items graded
n = 1,000 (0.1% each)
Best published result
75.6%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (75.6%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.

Snapshot September 8, 2026

Ranking on this test

57 models on SimpleQA Verified

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.1 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-6 AstraOpenAI75.6%Epoch-run
2Gemini 3.1 ProGoogle73.5%Epoch-run
3Claude Fable 5.1Anthropic70.8%Epoch-run
4Claude Fable 5Anthropic70.7%Epoch-run
5GPT-5.6 SolOpenAI69.7%Epoch-run
6Gemini 3.7 FlashGoogle69.2%Epoch-run
7Gemini 3 FlashGoogle66.8%Epoch-run
8Gemini 3.5 FlashGoogle66.2%Epoch-run
9Gemini 3.6 FlashGoogle66.2%Epoch-run
10GPT-5.5OpenAI63.0%Epoch-run
11Muse Spark 1.2Meta60.3%Epoch-run
12Claude Opus 5Anthropic59.9%Epoch-run
13Muse Spark 1.1Meta57.8%Epoch-run
14Qwen3.7 MaxAlibaba55.8%Epoch-run
15Claude Opus 4.8Anthropic53.0%Epoch-run
16DeepSeek V4 ProDeepSeek52.9%Epoch-run
17Qwen3.6 MaxAlibaba52.0%Epoch-run
18Claude Opus 4.7Anthropic51.7%Epoch-run
19Kimi K3Moonshot AI50.6%Epoch-run
20GPT-5OpenAI50.1%Epoch-run
21Grok 4.6xAI49.3%Epoch-run
22Qwen3 MaxAlibaba48.7%Epoch-run
23Grok 4.5xAI48.3%Epoch-run
24GPT-5.1OpenAI48.0%Epoch-run
25Claude Opus 4.6Anthropic47.0%Epoch-run
26GPT-5.4 ProOpenAI46.3%Epoch-run
27Qwen3.8 MaxAlibaba45.8%Epoch-run
28Claude Opus 4.5Anthropic45.7%Epoch-run
29GPT-5.4OpenAI45.1%Epoch-run
30Qwen3.6 PlusAlibaba44.1%Epoch-run
31GPT-5.6 TerraOpenAI43.2%Epoch-run
32GPT-5.6 LunaOpenAI41.0%Epoch-run
33GLM 5.3Zhipu AI41.0%Epoch-run
34InklingThinking Machines40.3%Epoch-run
35GPT-5.2OpenAI37.1%Epoch-run
36Kimi K2.7 CodeMoonshot AI36.5%Epoch-run
37Claude Sonnet 4.6Anthropic35.5%Epoch-run
38Kimi K2.6Moonshot AI34.9%Epoch-run
39Kimi K2.5Moonshot AI34.3%Epoch-run
40GLM 5.2Zhipu AI34.2%Epoch-run
41GLM 5.1Zhipu AI34.0%Epoch-run
42Claude Sonnet 5Anthropic33.7%Epoch-run
43DeepSeek V4 FlashDeepSeek33.6%Epoch-run
44Grok 4.3xAI33.2%Epoch-run
45GLM 4.7Zhipu AI32.2%Epoch-run
46Claude Sonnet 4.5Anthropic30.7%Epoch-run
47Grok 4.20xAI30.2%Epoch-run
48GPT-5.4 miniOpenAI29.4%Epoch-run
49Qwen3.5 PlusAlibaba25.4%Epoch-run
50GPT-5 miniOpenAI21.6%Epoch-run
51Qwen3.5 FlashAlibaba20.3%Epoch-run
52Inkling SmallThinking Machines19.1%Epoch-run
53Qwen3.6 FlashAlibaba15.9%Epoch-run
54Claude Haiku 4.5Anthropic13.2%Epoch-run
55GPT-5.4 nanoOpenAI11.7%Epoch-run
56GPT-5 nanoOpenAI11.7%Epoch-run
57Gemma 4 31BGoogle10.4%Epoch-run
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.