All benchmarks
Coding

CursorBench

Real coding-agent requests drawn from everyday editor sessions, scored against the change the developer actually shipped.

Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.

Read with care. Published by Cursor from its own product traffic and harness; not independently reproduced.

Publisher: Cursor

What a model is asked to do

Handle a real coding-agent request from an everyday editor session, judged against the change the developer actually shipped.

For exampleMake this module's configuration loading asynchronous and update every call site, keeping the behavior identical.

Why it matters. Judged against what a person shipped, so it measures usefulness inside a real workflow.

The example is original and illustrative, not an item from the dataset.

Models scored here
21
Who produced the numbers
Official leaderboard
Best published result
73.4%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (73.4%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

21 models on CursorBench

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Fable 5.1Anthropic73.4%Leaderboard
2Grok 4.6xAI70.8%Leaderboard
3Claude Fable 5Anthropic70.5%Leaderboard
4Claude Opus 5Anthropic70.0%Leaderboard
5Gemini 3.8 FlashGoogle69.2%Leaderboard
6GPT-5.6 SolOpenAI67.2%Leaderboard
7GPT-5.6 TerraOpenAI64.9%Leaderboard
8Claude Opus 4.7Anthropic64.8%Leaderboard
9Claude Opus 4.8Anthropic62.3%Leaderboard
10Gemini 3.7 FlashGoogle61.6%Leaderboard
11Claude Sonnet 5Anthropic61.5%Leaderboard
12GPT-5.6 LunaOpenAI61.1%Leaderboard
13Kimi K3Moonshot AI60.8%Leaderboard
14GPT-5.5OpenAI58.4%Leaderboard
15GLM 5.2Zhipu AI55.0%Leaderboard
16Gemini 3.6 FlashGoogle53.5%Leaderboard
17Gemini 3.5 FlashGoogle49.8%Leaderboard
18Kimi K2.7 CodeMoonshot AI49.7%Leaderboard
19Claude Sonnet 4.6Anthropic49.0%Leaderboard
20Kimi K2.6Moonshot AI47.6%Leaderboard
21Kimi K2.5Moonshot AI31.9%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.