SimpleQA Verified
A thousand verified short-answer factual questions — how often the model is right, without being allowed to search. Epoch AI runs it.
Recall of facts without hallucinating — short-answer accuracy. Knowledge is a single-test category in this snapshot (fewer than 2 comparable tests), so it is shown beside the Index and never averaged into it.
Publisher: Epoch AI (internal runs)What a model is asked to do
Answer a short factual question exactly, without searching, or say you do not know.
For exampleWhich architect designed the smaller second station that replaced the city's original 1901 harbor terminus?
Why it matters. How often a model is right about facts, and how rarely it invents one, decides whether you can trust it unsupervised.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 57
- Who produced the numbers
- Run by Epoch AI
- Items graded
- n = 1,000 (0.1% each)
- Best published result
- 75.6%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (75.6%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.
Snapshot September 8, 2026
Ranking on this test
57 models on SimpleQA Verified
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.1 points, so read gaps smaller than that as noise.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.