All benchmarks
ReasoningHard setInternally runnable

ARC-AGI-2

Abstract visual reasoning puzzles that are easy for people and designed to resist memorization — the benchmark built to measure general fluid intelligence. Semi-private evaluation set, verified by ARC Prize.

Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.

Read with care. Scores depend on the compute budget a lab chose; ARC Prize publishes cost per task beside every score and this ranking does not.

Publisher: ARC Prize Foundation

What a model is asked to do

Infer the hidden rule from a few input and output grid pairs, then apply it to a new grid.

For exampleThree examples show colored shapes being reflected across a diagonal and recolored by their size. Produce the output for a fourth, unseen input.

Why it matters. Novel visual puzzles resist memorization, so they measure fluid intelligence rather than recall.

The example is original and illustrative, not an item from the dataset.

Models scored here
51
Who produced the numbers
Official leaderboard
Items graded
n = 120 (0.8% each)
Best published result
95.0%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (95.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

51 models on ARC-AGI-2

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.8 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-6 AstraOpenAI95.0%Leaderboard
2GPT-5.6 SolOpenAI92.5%Leaderboard
3Claude Opus 5Anthropic90.4%Leaderboard
4Claude Fable 5.1Anthropic90.0%Leaderboard
5Claude Fable 5Anthropic89.2%Leaderboard
6GPT-5.5OpenAI85.0%Leaderboard
7Gemini 3.7 FlashGoogle84.6%Leaderboard
8GPT-5.5 ProOpenAI84.6%Leaderboard
9Gemini 3 Deep ThinkGoogle84.6%Leaderboard
10GPT-5.6 TerraOpenAI83.9%Leaderboard
11GPT-5.4 ProOpenAI83.3%Leaderboard
12Gemini 3.1 ProGoogle77.1%Leaderboard
13Claude Opus 4.7Anthropic75.8%Leaderboard
14GPT-5.4OpenAI74.0%Leaderboard
15Claude Opus 4.8Anthropic72.1%Leaderboard
16Gemini 3.5 FlashGoogle72.1%Leaderboard
17Claude Opus 4.6Anthropic69.2%Leaderboard
18Grok 4.6xAI67.1%Leaderboard
19Grok 4.20xAI65.1%Leaderboard
20DeepSeek V4 FlashDeepSeek61.4%Leaderboard
21DeepSeek V4 ProDeepSeek61.3%Leaderboard
22Kimi K3Moonshot AI60.4%Leaderboard
23Claude Sonnet 4.6Anthropic60.4%Leaderboard
24Gemini 3.6 FlashGoogle60.4%Leaderboard
25GPT-5.6 LunaOpenAI59.5%Leaderboard
26GPT-5.2 ProOpenAI54.2%Leaderboard
27GPT-5.2OpenAI52.9%Leaderboard
28Grok 4.5xAI52.6%Leaderboard
29Inkling SmallThinking Machines40.1%Leaderboard
30Claude Opus 4.5Anthropic37.6%Leaderboard
31InklingThinking Machines36.5%Leaderboard
32Gemini 3 FlashGoogle33.6%Leaderboard
33Gemini 3 ProGoogle31.1%Leaderboard
34GLM 5.2Zhipu AI22.8%Leaderboard
35GPT-5.4 miniOpenAI18.9%Leaderboard
36GPT-5 ProOpenAI18.3%Leaderboard
37GPT-5.1OpenAI17.6%Leaderboard
38Grok 4xAI16.0%Leaderboard
39Claude Sonnet 4.5Anthropic13.6%Leaderboard
40Kimi K2.5Moonshot AI11.8%Leaderboard
41Gemini 3.5 Flash-LiteGoogle10.3%Leaderboard
42GPT-5OpenAI9.9%Leaderboard
43GPT-5.4 nanoOpenAI5.7%Leaderboard
44Grok 4 FastxAI5.3%Leaderboard
45Gemini 2.5 ProGoogle4.9%Leaderboard
46GLM 5Zhipu AI4.9%Leaderboard
47MiniMax M2.5MiniMax4.9%Leaderboard
48GPT-5 miniOpenAI4.4%Leaderboard
49DeepSeek V3.2DeepSeek4.0%Leaderboard
50Claude Haiku 4.5Anthropic4.0%Leaderboard
51GPT-5 nanoOpenAI2.6%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.