All benchmarks
Coding

WeirdML

Unusual machine-learning tasks the model must solve end to end by writing and running training code.

Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.

Publisher: Håvard Ihle

What a model is asked to do

Solve an unusual machine-learning task end to end: write the training code, run it, and hit an accuracy target.

For exampleClassify the shapes in these noisy 32 by 32 images, where the label depends on rotation, using only the small training set provided.

Why it matters. Real machine-learning work is messy and unfamiliar. This rewards models that can experiment rather than recite a tutorial.

The example is original and illustrative, not an item from the dataset.

Models scored here
69
Who produced the numbers
Official leaderboard
Best published result
92.9%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (92.9%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

69 models on WeirdML

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1GPT-6 AstraOpenAI92.9%Leaderboard
2Claude Fable 5.1Anthropic92.9%Leaderboard
3Claude Fable 5Anthropic91.9%Leaderboard
4Claude Opus 5Anthropic91.8%Leaderboard
5GPT-5.6 Sol ProOpenAI89.4%Leaderboard
6GPT-5.6 SolOpenAI88.8%Leaderboard
7GPT-5.5OpenAI84.9%Leaderboard
8Claude Opus 4.8Anthropic82.9%Leaderboard
9Kimi K3Moonshot AI82.6%Leaderboard
10GPT-5.3 CodexOpenAI79.3%Leaderboard
11GPT-5.6 TerraOpenAI78.3%Leaderboard
12Claude Opus 4.6Anthropic78.0%Leaderboard
13GPT-5.4OpenAI77.7%Leaderboard
14Claude Opus 4.7Anthropic76.4%Leaderboard
15GLM 5.3Zhipu AI75.4%Leaderboard
16GPT-5.2OpenAI72.2%Leaderboard
17Gemini 3.1 ProGoogle72.1%Leaderboard
18GLM 5.2Zhipu AI70.1%Leaderboard
19Gemini 3 ProGoogle69.9%Leaderboard
20Claude Sonnet 5Anthropic68.8%Leaderboard
21Grok 4.6xAI67.3%Leaderboard
22DeepSeek V4 ProDeepSeek66.2%Leaderboard
23Claude Sonnet 4.6Anthropic66.1%Leaderboard
24Claude Opus 4.5Anthropic63.7%Leaderboard
25DeepSeek V4 FlashDeepSeek63.0%Leaderboard
26Gemini 3.5 FlashGoogle62.6%Leaderboard
27Gemini 3 FlashGoogle61.6%Leaderboard
28GPT-5.6 LunaOpenAI60.9%Leaderboard
29GPT-5.1OpenAI60.8%Leaderboard
30GPT-5OpenAI60.7%Leaderboard
31GPT-5 ProOpenAI60.4%Leaderboard
32GPT-5.4 miniOpenAI60.3%Leaderboard
33Muse Spark 1.2Meta60.3%Leaderboard
34GPT-5.4 ProOpenAI57.4%Leaderboard
35GLM 5.1Zhipu AI57.1%Leaderboard
36Gemini 3.6 FlashGoogle56.1%Leaderboard
37Kimi K2.6Moonshot AI55.9%Leaderboard
38GPT-5 CodexOpenAI54.5%Leaderboard
39Kimi K2.7 CodeMoonshot AI54.1%Leaderboard
40Gemini 2.5 ProGoogle54.0%Leaderboard
41GPT-5 miniOpenAI52.7%Leaderboard
42Grok 4.20xAI52.3%Leaderboard
43Gemma 4 31BGoogle52.3%Leaderboard
44Gemini 3.1 Flash-LiteGoogle52.2%Leaderboard
45Grok 4.3xAI49.9%Leaderboard
46GPT-5.4 nanoOpenAI49.2%Leaderboard
47GLM 5Zhipu AI48.2%Leaderboard
48GPT-OSS 120BOpenAI48.2%Leaderboard
49Claude Sonnet 4.5Anthropic47.7%Leaderboard
50DeepSeek V3.2DeepSeek46.7%Leaderboard
51Grok 4.5xAI46.4%Leaderboard
52Claude Opus 4.1Anthropic45.9%Leaderboard
53Grok 4xAI45.7%Leaderboard
54Kimi K2.5Moonshot AI45.6%Leaderboard
55Claude Haiku 4.5Anthropic45.4%Leaderboard
56Mistral Medium 3.5Mistral AI43.7%Leaderboard
57Nemotron 3 UltraNVIDIA43.5%Leaderboard
58Kimi K2Moonshot AI42.8%Leaderboard
59Grok 4 FastxAI42.9%Leaderboard
60Gemini 2.5 FlashGoogle41.9%Leaderboard
61GPT-OSS 20BOpenAI40.9%Leaderboard
62Qwen3.5 27BAlibaba39.5%Leaderboard
63Gemini 3.5 Flash-LiteGoogle39.0%Leaderboard
64GPT-5 nanoOpenAI38.1%Leaderboard
65Nemotron 3 SuperNVIDIA38.0%Leaderboard
66MiniMax M2.7MiniMax37.0%Leaderboard
67Gemma 4 26B A4BGoogle35.2%Leaderboard
68Qwen3.6 35B-A3BAlibaba34.5%Leaderboard
69InklingThinking Machines32.3%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.