Text Arena
Millions of blind head-to-head votes from real users. Style-controlled Arena score — the market's human-preference standard.
Blind head-to-head votes from real users on real prompts. Human preference is a single-test category in this snapshot (fewer than 2 comparable tests), so it is shown beside the Index and never averaged into it.
Publisher: ArenaWhat a model is asked to do
Answer the same prompt as a rival model; real users vote blind on which answer they prefer.
For exampleRewrite this dense lease clause for a first-time renter without losing the two conditions that actually matter.
Why it matters. Millions of blind votes on real prompts are the market's measure of what people prefer.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 92
- Who produced the numbers
- Official leaderboard
- Best published result
- 1,507
- License
- CC BY 4.0
- Direction
- Higher is better
- Scale
- Published as an Arena rating from blind head-to-head votes. For the Index, each rating becomes the expected win rate against the board leader under the Elo model, scaled so parity reads 100: 100 points behind the leader is a 36% win rate and reads 72. The scale does not depend on which weak model happens to be listed.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
92 models on Text Arena
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, as an expected win rate against the board leader). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 1,507 | 100.0 | board 2026-09-02 | Leaderboard |
| 2 | Claude Opus 4.6Anthropic | 1,505 | 99.3 | high · board 2026-09-02 | Leaderboard |
| 3 | Claude Fable 5.1Anthropic | 1,504 | 99.2 | max · board 2026-09-02 | Leaderboard |
| 4 | Claude Opus 4.7Anthropic | 1,502 | 98.5 | high · board 2026-09-02 | Leaderboard |
| 5 | Muse Spark 1.2Meta | 1,499 | 97.6 | xhigh · board 2026-09-02 | Leaderboard |
| 6 | Gemini 3.8 FlashGoogle | 1,494 | 96.2 | high · board 2026-09-02 | Leaderboard |
| 7 | Claude Opus 5Anthropic | 1,493 | 96.0 | high · board 2026-09-02 | Leaderboard |
| 8 | Muse Spark 1.1Meta | 1,492 | 95.7 | board 2026-09-02 | Leaderboard |
| 9 | Gemini 3.7 FlashGoogle | 1,491 | 95.3 | high · board 2026-09-02 | Leaderboard |
| 10 | Kimi K3Moonshot AI | 1,489 | 94.7 | max · board 2026-09-02 | Leaderboard |
| 11 | Muse SparkMeta | 1,488 | 94.5 | board 2026-09-02 | Leaderboard |
| 12 | Gemini 3.1 ProGoogle | 1,487 | 94.1 | preview · board 2026-09-02 | Leaderboard |
| 13 | Gemini 3 ProGoogle | 1,486 | 93.8 | board 2026-09-02 | Leaderboard |
| 14 | GPT-5.6 SolOpenAI | 1,483 | 93.1 | xhigh · board 2026-09-02 | Leaderboard |
| 15 | Claude Opus 4.8Anthropic | 1,482 | 92.9 | high · board 2026-09-02 | Leaderboard |
| 16 | GPT-5.5OpenAI | 1,482 | 92.8 | high · board 2026-09-02 | Leaderboard |
| 17 | GLM 5.3Zhipu AI | 1,482 | 92.8 | max · board 2026-09-02 | Leaderboard |
| 18 | Gemini 3.6 FlashGoogle | 1,480 | 92.3 | high · board 2026-09-02 | Leaderboard |
| 19 | Qwen3.8 MaxAlibaba | 1,480 | 92.1 | max · board 2026-09-02 | Leaderboard |
| 20 | Gemini 3.5 FlashGoogle | 1,479 | 92.0 | high · board 2026-09-02 | Leaderboard |
| 21 | GPT-5.4OpenAI | 1,477 | 91.2 | high · board 2026-09-02 | Leaderboard |
| 22 | Grok 4.20xAI | 1,475 | 90.6 | beta1 · board 2026-09-02 | Leaderboard |
| 23 | GLM 5.3 FlashZhipu AI | 1,474 | 90.5 | board 2026-09-02 | Leaderboard |
| 24 | Gemini 3 FlashGoogle | 1,474 | 90.4 | board 2026-09-02 | Leaderboard |
| 25 | Qwen3.7 MaxAlibaba | 1,474 | 90.4 | max-preview · board 2026-09-02 | Leaderboard |
| 26 | GPT-5.5 InstantOpenAI | 1,474 | 90.4 | board 2026-09-02 | Leaderboard |
| 27 | Claude Opus 4.5Anthropic | 1,473 | 90.2 | 20251101-high-32k · board 2026-09-02 | Leaderboard |
| 28 | Claude Sonnet 4.6Anthropic | 1,472 | 90.0 | board 2026-09-02 | Leaderboard |
| 29 | GLM 5.2Zhipu AI | 1,472 | 89.8 | max · board 2026-09-02 | Leaderboard |
| 30 | Grok 4.5xAI | 1,471 | 89.7 | board 2026-09-02 | Leaderboard |
| 31 | ERNIE 5.1Baidu | 1,468 | 88.8 | board 2026-09-02 | Leaderboard |
| 32 | MiMo V2.5 ProXiaomi | 1,468 | 88.8 | board 2026-09-02 | Leaderboard |
| 33 | GPT-5.6 TerraOpenAI | 1,466 | 88.3 | xhigh · board 2026-09-02 | Leaderboard |
| 34 | GLM 5.1Zhipu AI | 1,466 | 88.2 | board 2026-09-02 | Leaderboard |
| 35 | Grok 4.1xAI | 1,465 | 88.1 | thinking · board 2026-09-02 | Leaderboard |
| 36 | Claude Sonnet 5Anthropic | 1,462 | 87.2 | high · board 2026-09-02 | Leaderboard |
| 37 | Kimi K2.6Moonshot AI | 1,461 | 86.7 | board 2026-09-02 | Leaderboard |
| 38 | Qwen3.6 MaxAlibaba | 1,460 | 86.5 | max-preview · board 2026-09-02 | Leaderboard |
| 39 | DeepSeek V4 ProDeepSeek | 1,460 | 86.4 | high-20260813 · board 2026-09-02 | Leaderboard |
| 40 | GLM 5Zhipu AI | 1,458 | 85.9 | board 2026-09-02 | Leaderboard |
| 41 | Gemini 3.5 Flash-LiteGoogle | 1,457 | 85.6 | board 2026-09-02 | Leaderboard |
| 42 | Claude Sonnet 4.5Anthropic | 1,456 | 85.5 | 20250929-high-32k · board 2026-09-02 | Leaderboard |
| 43 | Seed 2.0 ProByteDance | 1,456 | 85.3 | board 2026-09-02 | Leaderboard |
| 44 | Hunyuan 3Tencent | 1,455 | 85.2 | board 2026-09-02 | Leaderboard |
| 45 | Qwen3.7 PlusAlibaba | 1,455 | 85.2 | board 2026-09-02 | Leaderboard |
| 46 | GPT-5.1OpenAI | 1,455 | 85.1 | high · board 2026-09-02 | Leaderboard |
| 47 | GPT-5.6 LunaOpenAI | 1,453 | 84.4 | xhigh · board 2026-09-02 | Leaderboard |
| 48 | Gemma 4 31BGoogle | 1,451 | 84.1 | board 2026-09-02 | Leaderboard |
| 49 | Kimi K2.5Moonshot AI | 1,451 | 83.9 | thinking · board 2026-09-02 | Leaderboard |
| 50 | Claude Opus 4.1Anthropic | 1,450 | 83.6 | 20250805-thinking-16k · board 2026-09-02 | Leaderboard |
| 51 | GPT-5.4 miniOpenAI | 1,448 | 83.2 | high · board 2026-09-02 | Leaderboard |
| 52 | Gemini 2.5 ProGoogle | 1,446 | 82.5 | board 2026-09-02 | Leaderboard |
| 53 | Qwen3.6 PlusAlibaba | 1,444 | 81.9 | board 2026-09-02 | Leaderboard |
| 54 | MiniMax M3MiniMax | 1,443 | 81.7 | board 2026-09-02 | Leaderboard |
| 55 | Grok 4.3xAI | 1,443 | 81.7 | board 2026-09-02 | Leaderboard |
| 56 | GLM 4.7Zhipu AI | 1,442 | 81.4 | board 2026-09-02 | Leaderboard |
| 57 | Qwen3.5 397B-A17BAlibaba | 1,441 | 81.2 | board 2026-09-02 | Leaderboard |
| 58 | InklingThinking Machines | 1,439 | 80.7 | board 2026-09-02 | Leaderboard |
| 59 | DeepSeek V4 FlashDeepSeek | 1,438 | 80.5 | high-preview · board 2026-09-02 | Leaderboard |
| 60 | Gemma 4 26B A4BGoogle | 1,438 | 80.5 | board 2026-09-02 | Leaderboard |
| 61 | GPT-5.2OpenAI | 1,438 | 80.3 | high · board 2026-09-02 | Leaderboard |
| 62 | Qwen3.8 27BAlibaba | 1,436 | 79.8 | board 2026-09-02 | Leaderboard |
| 63 | GPT-5OpenAI | 1,434 | 79.4 | high · board 2026-09-02 | Leaderboard |
| 64 | Qwen3 MaxAlibaba | 1,435 | 79.4 | max-preview · board 2026-09-02 | Leaderboard |
| 65 | MiMo V2.5Xiaomi | 1,434 | 79.2 | board 2026-09-02 | Leaderboard |
| 66 | Gemini 3.1 Flash-LiteGoogle | 1,432 | 78.8 | preview · board 2026-09-02 | Leaderboard |
| 67 | Kimi K2Moonshot AI | 1,430 | 78.2 | board 2026-09-02 | Leaderboard |
| 68 | Grok 4.1 FastxAI | 1,430 | 78.2 | reasoning · board 2026-09-02 | Leaderboard |
| 69 | Mistral Medium 3.5Mistral AI | 1,427 | 77.3 | board 2026-09-02 | Leaderboard |
| 70 | Nemotron 3 UltraNVIDIA | 1,426 | 77.1 | board 2026-09-02 | Leaderboard |
| 71 | DeepSeek V3.2DeepSeek | 1,425 | 76.8 | board 2026-09-02 | Leaderboard |
| 72 | GLM 4.6Zhipu AI | 1,425 | 76.7 | board 2026-09-02 | Leaderboard |
| 73 | Grok 4 FastxAI | 1,418 | 75.0 | board 2026-09-02 | Leaderboard |
| 74 | DeepSeek V3.1 TerminusDeepSeek | 1,418 | 74.8 | thinking · board 2026-09-02 | Leaderboard |
| 75 | Qwen3.5 122B-A10BAlibaba | 1,417 | 74.7 | board 2026-09-02 | Leaderboard |
| 76 | MiniMax M2.7MiniMax | 1,415 | 74.2 | board 2026-09-02 | Leaderboard |
| 77 | Mistral Large 3Mistral AI | 1,414 | 73.7 | board 2026-09-02 | Leaderboard |
| 78 | Claude Haiku 4.5Anthropic | 1,413 | 73.6 | 20251001 · board 2026-09-02 | Leaderboard |
| 79 | Grok 4xAI | 1,411 | 73.0 | board 2026-09-02 | Leaderboard |
| 80 | Gemini 2.5 FlashGoogle | 1,410 | 72.7 | board 2026-09-02 | Leaderboard |
| 81 | Qwen3.5 27BAlibaba | 1,408 | 72.2 | board 2026-09-02 | Leaderboard |
| 82 | Inkling SmallThinking Machines | 1,407 | 71.9 | board 2026-09-02 | Leaderboard |
| 83 | GPT-5.4 nanoOpenAI | 1,402 | 70.7 | high · board 2026-09-02 | Leaderboard |
| 84 | Qwen3.5 FlashAlibaba | 1,397 | 69.3 | board 2026-09-02 | Leaderboard |
| 85 | Qwen3.5 35B-A3BAlibaba | 1,395 | 68.8 | board 2026-09-02 | Leaderboard |
| 86 | MiniMax M2.5MiniMax | 1,391 | 67.6 | board 2026-09-02 | Leaderboard |
| 87 | GPT-5 miniOpenAI | 1,389 | 67.4 | high · board 2026-09-02 | Leaderboard |
| 88 | MiniMax M2.1MiniMax | 1,384 | 65.9 | preview · board 2026-09-02 | Leaderboard |
| 89 | GPT-OSS 120BOpenAI | 1,352 | 58.2 | board 2026-09-02 | Leaderboard |
| 90 | MiniMax M2MiniMax | 1,346 | 56.6 | board 2026-09-02 | Leaderboard |
| 91 | GPT-5 nanoOpenAI | 1,337 | 54.6 | high · board 2026-09-02 | Leaderboard |
| 92 | GPT-OSS 20BOpenAI | 1,317 | 50.2 | board 2026-09-02 | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.