All benchmarks
AgenticSelf-reported

Cybench

40 professional capture-the-flag cybersecurity challenges solved unguided. Share solved.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Read with care. Developer-reported: the numbers come from the labs' own model cards, collected by Epoch AI, not from an independent run.

Publisher: Stanford

What a model is asked to do

Solve professional capture-the-flag security challenges without guidance.

For exampleA web service leaks its session token through a timing side channel. Recover an admin session and read the flag.

Why it matters. Offensive security work demands systematic investigation and tool use under uncertainty.

The example is original and illustrative, not an item from the dataset.

Models scored here
7
Who produced the numbers
Developer-reported
Items graded
n = 40 (2.5% each)
Best published result
93.0%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (93.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Published by the model's own maker in a model card or launch post and collected by Epoch AI. Not an independent run.

Snapshot September 8, 2026

Ranking on this test

7 models on Cybench

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 2.5 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1Claude Opus 4.6Anthropic93.0%Self-reported
2Claude Opus 4.5Anthropic82.0%Self-reported
3Claude Sonnet 4.5Anthropic60.0%Self-reported
4Grok 4xAI43.0%Self-reported
5Claude Opus 4.1Anthropic42.0%Self-reported
6Grok 4.1xAI39.0%Self-reported
7Grok 4 FastxAI30.0%Self-reported
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.