All benchmarks
Agentic

APEX-Agents

Professional-services tasks — consulting, law, finance — completed as an agent and graded against expert rubrics. Pass@1.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Publisher: Mercor

What a model is asked to do

Complete professional-services work in consulting, law and finance as an agent, graded against expert rubrics.

For exampleReview the attached vendor contract, flag every clause that shifts liability to the client, and draft the redlines.

Why it matters. Expert-graded knowledge work is what professional teams would actually delegate.

The example is original and illustrative, not an item from the dataset.

Models scored here
49
Who produced the numbers
Official leaderboard
Best published result
47.4%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (47.4%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

49 models on APEX-Agents

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Fable 5.1Anthropic47.4%Leaderboard
2GPT-6 AstraOpenAI46.7%Leaderboard
3Claude Fable 5Anthropic45.0%Leaderboard
4Claude Opus 5Anthropic43.5%Leaderboard
5Claude Opus 4.8Anthropic42.5%Leaderboard
6Muse Spark 1.1Meta41.9%Leaderboard
7Grok 4.6xAI41.2%Leaderboard
8GPT-5.6 Sol ProOpenAI40.0%Leaderboard
9GPT-5.6 SolOpenAI39.9%Leaderboard
10Kimi K3Moonshot AI39.3%Leaderboard
11GPT-5.5OpenAI38.5%Leaderboard
12GPT-5.4OpenAI36.0%Leaderboard
13GLM 5.2Zhipu AI35.6%Leaderboard
14GPT-5.2OpenAI34.4%Leaderboard
15Grok 4.5xAI34.2%Leaderboard
16Claude Opus 4.7Anthropic33.9%Leaderboard
17Gemini 3.1 ProGoogle33.5%Leaderboard
18Claude Sonnet 5Anthropic32.5%Leaderboard
19Claude Opus 4.6Anthropic32.4%Leaderboard
20GPT-5.3 CodexOpenAI31.8%Leaderboard
21Gemini 3 ProGoogle31.5%Leaderboard
22Kimi K2.7 CodeMoonshot AI27.6%Leaderboard
23GPT-5.2 CodexOpenAI27.6%Leaderboard
24GPT-5.4 miniOpenAI24.6%Leaderboard
25Gemini 3 FlashGoogle24.0%Leaderboard
26Claude Sonnet 4.6Anthropic23.7%Leaderboard
27Claude Opus 4.5Anthropic20.7%Leaderboard
28GPT-5.1 CodexOpenAI20.7%Leaderboard
29GPT-5 CodexOpenAI20.1%Leaderboard
30Kimi K2.6Moonshot AI18.9%Leaderboard
31GPT-5OpenAI18.3%Leaderboard
32GPT-5.1OpenAI17.5%Leaderboard
33GLM 5Zhipu AI17.2%Leaderboard
34GPT-5.4 nanoOpenAI16.9%Leaderboard
35Grok 4xAI15.2%Leaderboard
36Kimi K2.5Moonshot AI14.4%Leaderboard
37Qwen3.5 PlusAlibaba13.6%Leaderboard
38Gemini 3.1 Flash-LiteGoogle13.0%Leaderboard
39Grok 4.1xAI12.8%Leaderboard
40Nemotron 3 UltraNVIDIA11.5%Leaderboard
41Claude Haiku 4.5Anthropic8.9%Leaderboard
42GLM 4.7Zhipu AI8.7%Leaderboard
43DeepSeek V3.2DeepSeek7.0%Leaderboard
44Gemini 2.5 ProGoogle6.6%Leaderboard
45MiniMax M2.5MiniMax6.2%Leaderboard
46GPT-OSS 120BOpenAI4.7%Leaderboard
47Kimi K2Moonshot AI4.1%Leaderboard
48GLM 4.6Zhipu AI4.0%Leaderboard
49Gemini 2.5 FlashGoogle1.8%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.