Benchmarks
A free, public ranking of every notable AI model on the market. Each model's scores on third-party benchmarks are measured against the best published result on each test, averaged by category, and combined into one OpenCharts Index with a letter grade and a rank band. No sign-in needed: open opencharts.com/benchmarks.
What it is
OpenCharts publishes no scores of its own for the ranked models. It reads what publishers other than the models' makers publish (Epoch AI's Benchmarking Hub and the Arena leaderboard dataset, both Creative Commons Attribution), keeps the best published run per model and test, measures every score against the best published result on that test, and averages. Every number on the site links back to the publisher that produced it and says who produced it, and the snapshot date is printed on every page. Every model also carries its lab's country of origin.
The hub reads 31 benchmarks across 8 categories: Reasoning, Knowledge, Coding, Math, Agentic, Multimodal, Human preference, Long context. Only categories with at least 2 comparable tests are averaged into the Index; a category that is a single benchmark is shown beside it.
Reading the leaderboard
Search, filter, sort
Type a model, lab or country; tap lab logos or country flags to filter; pick a category; or toggle open weights, third-party published only, or Available in Theo. Click any column header to sort by it (click again to flip), or use the Sort menu. The URL updates as you go, so a filtered view is a link you can share.
Grade, Index and rank band
The colored pill is the letter grade; the number beside it is the OpenCharts Index (0 to 100). Under the rank sits its band: the places the model could hold if any single test were dropped. Neighbours whose bands overlap are ties. Hover a grade for what its band means and what the number rests on. A grey grade means the model is provisional: shown, but not ranked yet.
Coverage and the full table
The Tests column shows how many comparable tests and counting categories a model has. Switch to Full table to see every column: lab, country of origin, every category mean (single-test categories ringed, because they sit beside the Index), the hard-set score, release date, weights and Theo availability.
Who produced each number
A green pill on a model means every counted score was published by a third party. Beside each score a smaller label says which kind: run by Epoch AI, from an official leaderboard or third-party evaluator, or developer-reported (amber). An amber model pill means OpenCharts ran the tests itself and no publisher has verified the engine yet.
The OpenCharts Index and grades
- The best published run per model and benchmark counts.
- Each score is measured against the frontier on its test: the best result the publisher lists across its whole file, whether or not that model is tracked here. Percent scores become the share of that frontier above random guessing, so the leader reads 100 and sitting a harder exam never lowers a model. Arena ratings become the expected win rate against the board leader (100 points behind reads 72). Open-ended values sit on a log scale above a fixed floor. A test with fewer than 3 models scored is shown but not counted.
- Scores are averaged inside each category. A category is part of the Index once the snapshot has at least 2 comparable tests in it, and a model's mean counts once it has results on at least half of those tests.
- The OpenCharts Index is the mean of the counting category means, so a lab cannot climb by publishing many tests of one kind.
- A model is ranked once it has at least 6 comparable tests across 3 counting categories; otherwise it is provisional. The grade reads the Index as a distance from the frontier: A+ at 95+, A at 90+, A- at 85+, B+ at 80+, B at 75+, B- at 70+, C+ at 65+, C at 60+, C- at 50+, D at 40+, F at 0+. Within 5% of the best published results, on average.
- Every Index carries a leave-one-out band (where it lands if any single test is dropped) and every rank carries the places that band covers. Head-to-head comparisons call any gap under 2 points a tie, because one item on a 45-problem exam is 2.2 points.
Category leaders are the third-party published, ranked models with the best qualifying mean in each category; leaders of single-test categories are marked as sitting beside the Index. The full method, with every threshold, is on the methodology page.
Provenance and verification labels
Third-party published means a publisher other than the model's maker produced every counted score. It is not the same claim as “independently re-run”, so every score also carries one of three provenance labels, in Epoch AI's own classification of its hub: Epoch-run (Epoch AI ran the evaluation itself, under settings it documents per model, with a public log), Leaderboard (the benchmark's official leaderboard or a third-party evaluator), and Self-reported (the model's maker published the number in a model card or launch post and Epoch AI collected it). One benchmark is developer-reported throughout (Cybench); their cards and pages say so in amber.
Internal, independent verification TBD means no publisher has scored the engine yet, and the scores shown are OpenCharts' own runs on the same public tests: each card shows the dataset version, item count, repeats, grader and run date. Internal scores rank in the same table but never lead a spotlight, and the label stays until a publisher verifies the engine. Tests OpenCharts has not run for such an engine read TBD rather than a dash.
The hard-set score
The hard-set score is the mean of a model's relative scores on the 8 hardest, most general evaluations in the set (Humanity's Last Exam, ARC-AGI-2, SWE-bench Verified, FrontierMath (Tiers 1–3), FrontierMath Tier 4, Terminal-Bench, GDPval, METR Time Horizon), shown once a model has published results on at least 4 of them and counting third-party scores only. A newly released model usually has fewer; it is listed as waiting on more results, with the mean it has so far, and enters the ranking as publishers add results.
Known limits
- The best published run counts, so labs that publish many effort levels get the benefit of their best one.
- Agent benchmarks compare model-plus-harness systems; harness choice can move a score by ten points.
- Some benchmarks are authored or graded by a lab that also competes on them. Every such test page says so.
- Publishers' confidence intervals are not propagated; the leave-one-out band is a stand-in.
- The frontier moves whenever a better result is published, so Index values compare within a snapshot, not across snapshots.
Models Theo runs on

Models that Theo orchestrates carry an Available in Theo mark on the leaderboard and a card on their page naming the Theo tier they serve. Theo picks the right engine for each step and always shows which one answered; the Theo Engines article explains how. The Available in Theo filter shows only those models.
For agents and spreadsheets
The whole ranking is published from the same data the pages render:
- /benchmarks/llms.txt as plain text, in the llms.txt style AI assistants read.
- /benchmarks/rankings.json as JSON with every model, rank band, category mean, score, frontier, source and provenance.
Credit the publishers as their licenses require and OpenCharts for the Index; the attribution lines are on the methodology page.
FAQ
How often is the ranking updated?
The snapshot is refreshed on a weekly schedule and whenever a notable model ships. The capture date is printed on every page and in the machine-readable files.
Why is a model I know of missing?
A model appears once at least one publisher has scored it and the registry knows its name. Newly released models usually land within a week of their first published result.
Why does the Index differ from a publisher's own composite?
Each publisher weights and normalizes differently. The OpenCharts Index is one transparent rule applied to all of them: measure against the frontier, average by category, average the categories. The methodology page spells it out.
Two models are one place apart. Is one better?
Not necessarily. Read the rank band under each rank: if the two bands overlap, the order between them is not meaningful and the site treats them as a tie. Head-to-head comparisons on a model page call any gap under two points even.
Do I need an account?
No. Every benchmarks page and both machine-readable files are free and public.
Open Theo Chat and ask. Theo knows how the ranking is built and points you to the right page; inside a chat, it always shows which engine answered.Related Articles
Theo Engines
Every Theo-branded engine powering the product — generated directly from the unified model catalog.
AI Chat
Conversational AI that creates flowcharts, whiteboards, notes, and presentations from natural language.
Consensus Mode
Query 7 AI models simultaneously, see agreement levels, and get a synthesized consensus answer.
Theo AI Overview
Overview of Theo AI — 11 chat modes, creation tools, generative UI, memory, and the sidebar mascot.