Vending-Bench 2
Run a simulated vending business for a year from a $500 float — ordering, pricing, cash flow. Mean final balance over five runs; normalized on a log scale from the starting balance to the frontier.
Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.
Read with care. Andon Labs estimates a strong human operator at roughly $63,000, so every model is far from the ceiling.
Publisher: Andon LabsWhat a model is asked to do
Run a simulated vending business for a year: ordering, pricing and cash flow. The score is the final balance.
For exampleStart with a small cash float and one machine; negotiate supplier prices, set a menu and keep it stocked through the seasons.
Why it matters. Long-horizon decisions whose consequences compound. It exposes models that drift or forget.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 53
- Who produced the numbers
- Official leaderboard
- Best published result
- $11,182
- Log-scale floor
- $500
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- An open-ended value. For the Index, each score is placed on a log scale from a fixed floor ($500) to the best published result ($11,182), so the frontier reads 100 and doubling counts the same anywhere on the scale.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
53 models on Vending-Bench 2
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, on a log scale). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.