All benchmarks
Reasoning

SimpleBench

Trick questions about the everyday world where humans score in the eighties — spatio-temporal reasoning, social intelligence and adversarial phrasing.

Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.

Publisher: SimpleBench / LM Council

What a model is asked to do

Answer everyday trick questions about the physical and social world that most people find easy.

For exampleSofia puts an ice cube in a hot pan and leaves the kitchen for an hour. When she comes back and looks in the pan, what does she see?

Why it matters. Common sense about the real world is where fluent models still trip. People score in the eighties here.

The example is original and illustrative, not an item from the dataset.

Models scored here
46
Who produced the numbers
Official leaderboard
Best published result
81.9%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (81.9%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

46 models on SimpleBench

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Fable 5Anthropic81.9%Leaderboard
2Claude Opus 5Anthropic80.6%Leaderboard
3Gemini 3.1 ProGoogle79.6%Leaderboard
4GPT-5.5 ProOpenAI76.9%Leaderboard
5Gemini 3.5 FlashGoogle76.7%Leaderboard
6Gemini 3 ProGoogle76.4%Leaderboard
7Grok 4.6xAI75.9%Leaderboard
8Muse Spark 1.2Meta74.5%Leaderboard
9GPT-5.4 ProOpenAI74.1%Leaderboard
10GPT-5.6 Sol ProOpenAI71.7%Leaderboard
11Qwen3.7 MaxAlibaba70.4%Leaderboard
12Grok 4.5xAI70.0%Leaderboard
13GPT-5.5OpenAI69.0%Leaderboard
14Claude Opus 4.6Anthropic67.6%Leaderboard
15GPT-5.6 SolOpenAI64.8%Leaderboard
16Claude Opus 4.8Anthropic64.8%Leaderboard
17Qwen3.6 MaxAlibaba63.0%Leaderboard
18Claude Opus 4.5Anthropic62.0%Leaderboard
19Claude Opus 4.7Anthropic61.7%Leaderboard
20GPT-5 ProOpenAI61.6%Leaderboard
21DeepSeek V4 FlashDeepSeek61.1%Leaderboard
22Gemini 3 FlashGoogle61.1%Leaderboard
23Kimi K3Moonshot AI60.7%Leaderboard
24Claude Sonnet 5Anthropic60.6%Leaderboard
25Grok 4xAI60.5%Leaderboard
26Claude Opus 4.1Anthropic60.0%Leaderboard
27GLM 5.2Zhipu AI58.8%Leaderboard
28Kimi K2.7 CodeMoonshot AI57.9%Leaderboard
29GPT-5.2 ProOpenAI57.4%Leaderboard
30GPT-5OpenAI56.7%Leaderboard
31Grok 4.1 FastxAI56.0%Leaderboard
32GLM 5.1Zhipu AI55.1%Leaderboard
33Claude Sonnet 4.5Anthropic54.3%Leaderboard
34GPT-5.1OpenAI53.2%Leaderboard
35GLM 5Zhipu AI53.2%Leaderboard
36DeepSeek V3.2DeepSeek52.6%Leaderboard
37InklingThinking Machines50.0%Leaderboard
38GPT-5.6 TerraOpenAI48.9%Leaderboard
39GLM 4.7Zhipu AI47.7%Leaderboard
40GPT-5.6 LunaOpenAI46.8%Leaderboard
41Kimi K2.5Moonshot AI46.8%Leaderboard
42GPT-5.2OpenAI45.8%Leaderboard
43MiniMax M3MiniMax45.8%Leaderboard
44Gemini 2.5 FlashGoogle41.2%Leaderboard
45Qwen3.6 FlashAlibaba35.2%Leaderboard
46GPT-OSS 120BOpenAI22.1%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.