All benchmarks
MultimodalBeside the Index

Vision Arena

Blind head-to-head votes on prompts that include an image. Style-controlled Arena score.

Understanding images alongside text, judged by human preference. Multimodal is a single-test category in this snapshot (fewer than 2 comparable tests), so it is shown beside the Index and never averaged into it.

Publisher: Arena

What a model is asked to do

Answer prompts that include an image; real users vote blind between two models' answers.

For exampleHere is a photo of a circuit board after a power surge. Which component most likely failed, and what points to it?

Why it matters. Seeing is half of most real tasks: screenshots, documents, photos and charts.

The example is original and illustrative, not an item from the dataset.

Models scored here
54
Who produced the numbers
Official leaderboard
Best published result
1,313
License
CC BY 4.0
Direction
Higher is better
Scale
Published as an Arena rating from blind head-to-head votes. For the Index, each rating becomes the expected win rate against the board leader under the Elo model, scaled so parity reads 100: 100 points behind the leader is a 36% win rate and reads 72. The scale does not depend on which weak model happens to be listed.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

54 models on Vision Arena

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, as an expected win rate against the board leader). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Fable 5Anthropic1,313Leaderboard
2Claude Opus 4.7Anthropic1,301Leaderboard
3Qwen3.8 MaxAlibaba1,300Leaderboard
4Claude Opus 4.6Anthropic1,299Leaderboard
5Muse SparkMeta1,294Leaderboard
6Muse Spark 1.2Meta1,292Leaderboard
7Claude Opus 5Anthropic1,290Leaderboard
8Gemini 3 ProGoogle1,289Leaderboard
9GPT-5.5OpenAI1,286Leaderboard
10Claude Opus 4.8Anthropic1,285Leaderboard
11Gemini 3.6 FlashGoogle1,285Leaderboard
12Gemini 3.5 FlashGoogle1,284Leaderboard
13GPT-5.6 SolOpenAI1,282Leaderboard
14Grok 4.5xAI1,282Leaderboard
15GPT-5.4OpenAI1,281Leaderboard
16Muse Spark 1.1Meta1,279Leaderboard
17Gemini 3.1 ProGoogle1,278Leaderboard
18GPT-5.5 InstantOpenAI1,278Leaderboard
19Claude Sonnet 4.6Anthropic1,275Leaderboard
20Gemini 3 FlashGoogle1,272Leaderboard
21GLM 5.3 FlashZhipu AI1,273Leaderboard
22Claude Sonnet 5Anthropic1,267Leaderboard
23Gemini 3.5 Flash-LiteGoogle1,266Leaderboard
24GPT-5.6 TerraOpenAI1,266Leaderboard
25Qwen3.7 PlusAlibaba1,266Leaderboard
26Kimi K2.6Moonshot AI1,263Leaderboard
27Gemma 4 31BGoogle1,261Leaderboard
28Seed 2.0 ProByteDance1,257Leaderboard
29Grok 4.20xAI1,256Leaderboard
30GPT-5.6 LunaOpenAI1,254Leaderboard
31GPT-5.4 miniOpenAI1,252Leaderboard
32Qwen3.8 27BAlibaba1,251Leaderboard
33GPT-5.1OpenAI1,250Leaderboard
34Kimi K2.5Moonshot AI1,250Leaderboard
35Qwen3.5 397B-A17BAlibaba1,247Leaderboard
36Gemini 2.5 ProGoogle1,246Leaderboard
37GPT-5.2OpenAI1,244Leaderboard
38Gemma 4 26B A4BGoogle1,242Leaderboard
39Grok 4.3xAI1,241Leaderboard
40MiniMax M3MiniMax1,237Leaderboard
41MiMo V2.5Xiaomi1,235Leaderboard
42Gemini 3.1 Flash-LiteGoogle1,235Leaderboard
43Qwen3.5 122B-A10BAlibaba1,227Leaderboard
44Qwen3.5 27BAlibaba1,219Leaderboard
45Gemini 2.5 FlashGoogle1,214Leaderboard
46GPT-5OpenAI1,210Leaderboard
47Inkling SmallThinking Machines1,207Leaderboard
48Mistral Large 3Mistral AI1,205Leaderboard
49GPT-5.4 nanoOpenAI1,201Leaderboard
50Mistral Medium 3.5Mistral AI1,199Leaderboard
51Grok 4.1 FastxAI1,195Leaderboard
52Grok 4xAI1,184Leaderboard
53GPT-5 miniOpenAI1,182Leaderboard
54GPT-5 nanoOpenAI1,145Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.