FrontierCode
Hard, realistic software tasks run through a coding-agent harness. Main score.
Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.
Read with care. Published by Cognition, a coding-agent vendor, from its own harness; not independently reproduced.
Publisher: CognitionWhat a model is asked to do
Complete a hard, realistic software task through a coding-agent harness.
For exampleAdd end-to-end encryption to the attachment path of this messaging service, including key rotation, with the existing tests still passing.
Why it matters. Multi-file, multi-hour engineering tasks are where agentic coding is heading.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 27
- Who produced the numbers
- Official leaderboard
- Best published result
- 53.5%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (53.5%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
27 models on FrontierCode
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 53.5% | 100.0 | harness: claude-code · reasoning effort: xhigh | Leaderboard |
| 2 | Claude Opus 5Anthropic | 53.4% | 99.8 | max · harness: claude-code · reasoning effort: medium | Leaderboard |
| 3 | GPT-6 AstraOpenAI | 53.3% | 99.6 | max · harness: codex · reasoning effort: max | Leaderboard |
| 4 | Claude Fable 5.1Anthropic | 50.9% | 95.2 | medium · harness: claude-code · reasoning effort: medium | Leaderboard |
| 5 | Grok 4.6xAI | 48.0% | 89.8 | harness: grok-build · reasoning effort: high | Leaderboard |
| 6 | GPT-5.6 SolOpenAI | 47.5% | 88.8 | harness: codex · reasoning effort: max | Leaderboard |
| 7 | Claude Opus 4.8Anthropic | 46.5% | 86.9 | harness: claude-code · reasoning effort: max | Leaderboard |
| 8 | Kimi K3Moonshot AI | 44.2% | 82.6 | harness: mini-swe-agent · reasoning effort: none | Leaderboard |
| 9 | Gemini 3.7 FlashGoogle | 43.6% | 81.5 | harness: chisel · reasoning effort: medium | Leaderboard |
| 10 | GPT-5.5OpenAI | 43.0% | 80.3 | harness: codex · reasoning effort: xhigh | Leaderboard |
| 11 | Claude Sonnet 5Anthropic | 42.7% | 79.9 | harness: claude-code · reasoning effort: xhigh | Leaderboard |
| 12 | Grok 4.5xAI | 42.4% | 79.4 | harness: grok-build · reasoning effort: high | Leaderboard |
| 13 | GPT-5.6 TerraOpenAI | 41.3% | 77.2 | harness: codex · reasoning effort: max | Leaderboard |
| 14 | GPT-5.6 LunaOpenAI | 39.8% | 74.4 | harness: codex · reasoning effort: max | Leaderboard |
| 15 | Claude Opus 4.7Anthropic | 38.5% | 72.1 | harness: claude-code · reasoning effort: max | Leaderboard |
| 16 | Gemini 3.6 FlashGoogle | 34.4% | 64.3 | harness: chisel · reasoning effort: medium | Leaderboard |
| 17 | Kimi K2.7 CodeMoonshot AI | 30.1% | 56.2 | harness: mini-swe-agent · reasoning effort: none | Leaderboard |
| 18 | GPT-5.4 miniOpenAI | 27.0% | 50.6 | 2026-03-17 · harness: codex · reasoning effort: xhigh | Leaderboard |
| 19 | Claude Opus 4.6Anthropic | 26.6% | 49.8 | harness: claude-code · reasoning effort: high | Leaderboard |
| 20 | GLM 5.2Zhipu AI | 24.5% | 45.8 | none · harness: mini-swe-agent · reasoning effort: none | Leaderboard |
| 21 | Claude Sonnet 4.6Anthropic | 24.3% | 45.5 | harness: claude-code · reasoning effort: max | Leaderboard |
| 22 | DeepSeek V4 FlashDeepSeek | 18.8% | 35.2 | harness: chisel · reasoning effort: high | Leaderboard |
| 23 | DeepSeek V4 ProDeepSeek | 17.6% | 33.0 | none · harness: mini-swe-agent · reasoning effort: none | Leaderboard |
| 24 | MiniMax M3MiniMax | 14.7% | 27.5 | harness: mini-swe-agent · reasoning effort: none | Leaderboard |
| 25 | InklingThinking Machines | 14.0% | 26.2 | harness: mini-swe-agent · reasoning effort: 0.99 | Leaderboard |
| 26 | Qwen3.7 PlusAlibaba | 10.2% | 19.1 | harness: mini-swe-agent · reasoning effort: none | Leaderboard |
| 27 | Mistral Medium 3.5Mistral AI | 8.0% | 15.0 | harness: chisel · reasoning effort: none | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.