All benchmarks
Agentic

DeepResearch Bench

PhD-level research briefs written by an agent with web access, graded on coverage, insight and citations.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Publisher: DeepResearch Bench

What a model is asked to do

Write a PhD-level research brief with web access, graded on coverage, insight and citations.

For exampleSurvey the last three years of room-temperature superconductivity claims and assess which replications hold up.

Why it matters. A research agent lives or dies on whether its sources are real and its synthesis is right.

The example is original and illustrative, not an item from the dataset.

Models scored here
19
Who produced the numbers
Official leaderboard
Items graded
n = 100 (1.0% each)
Best published result
55.3%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (55.3%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

19 models on DeepResearch Bench

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 1.0 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1Claude Opus 4.6Anthropic55.3%Leaderboard
2Claude Sonnet 4.6Anthropic54.9%Leaderboard
3Claude Opus 4.5Anthropic54.8%Leaderboard
4GPT-5.5OpenAI54.0%Leaderboard
5Claude Sonnet 4.5Anthropic52.6%Leaderboard
6Claude Opus 4.8Anthropic50.2%Leaderboard
7Gemini 3 FlashGoogle49.8%Leaderboard
8GPT-5OpenAI49.6%Leaderboard
9Claude Opus 4.1Anthropic48.3%Leaderboard
10Gemini 3.1 ProGoogle47.8%Leaderboard
11Grok 4xAI47.3%Leaderboard
12Gemini 3 ProGoogle46.3%Leaderboard
13Claude Haiku 4.5Anthropic45.5%Leaderboard
14GPT-5.1OpenAI42.8%Leaderboard
15Gemini 2.5 ProGoogle41.5%Leaderboard
16GPT-5.2OpenAI41.1%Leaderboard
17Gemini 3.1 Flash-LiteGoogle37.3%Leaderboard
18GPT-5.4 miniOpenAI36.3%Leaderboard
19GPT-5.4OpenAI35.1%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.