All benchmarks
Coding

Code Arena (WebDev)

Blind head-to-head votes on web apps two models built from the same prompt. Arena score.

Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.

Publisher: Arena

What a model is asked to do

Build a web app from the same prompt as a rival model; real users vote blind on the result.

For exampleBuild a kanban board with drag-and-drop, a dark mode and local persistence, all in a single page.

Why it matters. People judging finished apps side by side is the most honest measure of front-end quality.

The example is original and illustrative, not an item from the dataset.

Models scored here
86
Who produced the numbers
Official leaderboard
Best published result
1,797
License
CC BY 4.0
Direction
Higher is better
Scale
Published as an Arena rating from blind head-to-head votes. For the Index, each rating becomes the expected win rate against the board leader under the Elo model, scaled so parity reads 100: 100 points behind the leader is a 36% win rate and reads 72. The scale does not depend on which weak model happens to be listed.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

86 models on Code Arena (WebDev)

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, as an expected win rate against the board leader). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1GPT-6 AstraOpenAI1,797Leaderboard
2Claude Fable 5.1Anthropic1,762Leaderboard
3Claude Opus 5Anthropic1,688Leaderboard
4Qwen3.8 MaxAlibaba1,686Leaderboard
5Kimi K3Moonshot AI1,674Leaderboard
6Claude Fable 5Anthropic1,629Leaderboard
7Qwen3.8 Flash NextAlibaba1,626Leaderboard
8Grok 4.6xAI1,625Leaderboard
9Muse Spark 1.3Meta1,622Leaderboard
10GPT-5.6 SolOpenAI1,617Leaderboard
11GLM 5.3Zhipu AI1,609Leaderboard
12GLM 5.3 FlashZhipu AI1,605Leaderboard
13Qwen3.8 27BAlibaba1,594Leaderboard
14Gemini 3.7 FlashGoogle1,587Leaderboard
15GLM 5.2Zhipu AI1,587Leaderboard
16DeepSeek V4 ProDeepSeek1,582Leaderboard
17DeepSeek V4 FlashDeepSeek1,580Leaderboard
18Gemini 3.8 FlashGoogle1,567Leaderboard
19Claude Opus 4.8Anthropic1,563Leaderboard
20Claude Opus 4.7Anthropic1,557Leaderboard
21Grok 4.5xAI1,556Leaderboard
22Claude Opus 4.6Anthropic1,546Leaderboard
23Muse Spark 1.1Meta1,541Leaderboard
24Gemini 3.6 FlashGoogle1,538Leaderboard
25Claude Sonnet 5Anthropic1,537Leaderboard
26Muse Spark 1.2Meta1,534Leaderboard
27Claude Sonnet 4.6Anthropic1,521Leaderboard
28GPT-5.6 TerraOpenAI1,520Leaderboard
29GPT-5.6 LunaOpenAI1,519Leaderboard
30Qwen3.7 MaxAlibaba1,517Leaderboard
31Hunyuan 3Tencent1,512Leaderboard
32GPT-5.5OpenAI1,510Leaderboard
33Kimi K2.6Moonshot AI1,509Leaderboard
34GLM 5.1Zhipu AI1,508Leaderboard
35Gemini 3.5 FlashGoogle1,500Leaderboard
36Claude Opus 4.5Anthropic1,495Leaderboard
37MiniMax M3MiniMax1,487Leaderboard
38Qwen3.6 MaxAlibaba1,479Leaderboard
39MiMo V2.5 ProXiaomi1,475Leaderboard
40Kimi K2.7 CodeMoonshot AI1,472Leaderboard
41GPT-5.4OpenAI1,463Leaderboard
42Qwen3.6 PlusAlibaba1,460Leaderboard
43Gemini 3.5 Flash-LiteGoogle1,449Leaderboard
44Gemini 3.1 ProGoogle1,446Leaderboard
45Gemini 3 ProGoogle1,439Leaderboard
46Gemini 3 FlashGoogle1,438Leaderboard
47MiMo V2.5Xiaomi1,438Leaderboard
48Kimi K2.5Moonshot AI1,436Leaderboard
49GLM 5Zhipu AI1,436Leaderboard
50GLM 4.7Zhipu AI1,434Leaderboard
51GPT-5OpenAI1,419Leaderboard
52GPT-5.2OpenAI1,416Leaderboard
53InklingThinking Machines1,409Leaderboard
54GPT-5.3 CodexOpenAI1,409Leaderboard
55Inkling SmallThinking Machines1,405Leaderboard
56Qwen3.5 397B-A17BAlibaba1,399Leaderboard
57MiniMax M2.7MiniMax1,398Leaderboard
58GPT-5.4 miniOpenAI1,397Leaderboard
59Claude Sonnet 4.5Anthropic1,392Leaderboard
60GPT-5.1OpenAI1,392Leaderboard
61Claude Opus 4.1Anthropic1,389Leaderboard
62MiniMax M2.1MiniMax1,387Leaderboard
63MiniMax M2.5MiniMax1,384Leaderboard
64Grok 4.20xAI1,374Leaderboard
65Gemma 4 31BGoogle1,363Leaderboard
66Gemma 4 26B A4BGoogle1,361Leaderboard
67DeepSeek V3.2DeepSeek1,360Leaderboard
68Qwen3.5 122B-A10BAlibaba1,358Leaderboard
69Qwen3.5 27BAlibaba1,357Leaderboard
70Grok 4.3xAI1,356Leaderboard
71GLM 4.6Zhipu AI1,340Leaderboard
72GPT-5.2 CodexOpenAI1,338Leaderboard
73GPT-5.1 CodexOpenAI1,336Leaderboard
74Claude Haiku 4.5Anthropic1,329Leaderboard
75Kimi K2Moonshot AI1,322Leaderboard
76MiniMax M2MiniMax1,298Leaderboard
77Mistral Medium 3.5Mistral AI1,265Leaderboard
78Gemini 3.1 Flash-LiteGoogle1,254Leaderboard
79Qwen3.5 35B-A3BAlibaba1,250Leaderboard
80GPT-5.1 Codex MiniOpenAI1,244Leaderboard
81Grok 4.1 FastxAI1,240Leaderboard
82Qwen3.5 FlashAlibaba1,238Leaderboard
83Mistral Large 3Mistral AI1,229Leaderboard
84Gemini 2.5 ProGoogle1,226Leaderboard
85Grok 4.1xAI1,211Leaderboard
86Grok 4 FastxAI1,162Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.