All benchmarks
AgenticHard set

METR Time Horizon

The length of software task (in human-expert minutes) a model completes with 50% reliability. Shown in minutes; normalized on a log scale from one minute to the frontier.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Read with care. METR publishes wide confidence intervals around each horizon; the point estimate is used here.

Publisher: METR

What a model is asked to do

Complete software tasks of increasing length; the score is the task length a model finishes with 50% reliability.

For exampleTasks run from a five-minute regex fix to a multi-hour feature build with its own test plan.

Why it matters. The horizon a model can work through unsupervised is the clearest measure of how much you can delegate.

The example is original and illustrative, not an item from the dataset.

Models scored here
14
Who produced the numbers
Official leaderboard
Best published result
17.4 h
Log-scale floor
1 min
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
An open-ended value. For the Index, each score is placed on a log scale from a fixed floor (1 min) to the best published result (17.4 h), so the frontier reads 100 and doubling counts the same anywhere on the scale.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

14 models on METR Time Horizon

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, on a log scale). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Opus 4.6Anthropic12.0 hLeaderboard
2Gemini 3.1 ProGoogle6.4 hLeaderboard
3GPT-5.2OpenAI5.9 hLeaderboard
4GPT-5.3 CodexOpenAI5.8 hLeaderboard
5GPT-5.4OpenAI5.7 hLeaderboard
6Claude Opus 4.5Anthropic4.9 hLeaderboard
7Gemini 3 ProGoogle3.7 hLeaderboard
8GPT-5.1 CodexOpenAI3.7 hLeaderboard
9GPT-5OpenAI3.4 hLeaderboard
10Claude Sonnet 4.5Anthropic2.0 hLeaderboard
11Claude Opus 4.1Anthropic114 minLeaderboard
12Grok 4xAI110 minLeaderboard
13Kimi K2Moonshot AI54 minLeaderboard
14GPT-OSS 120BOpenAI42 minLeaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.