Terminal-Bench
https://www.tbench.ai/leaderboard
- Kind
- benchmark
- Maintainer
- Stanford / Laude Institute
- Feeds axes
- agentic
- Measures
- Whether a model can complete real multi-step technical tasks on its own — fixing a broken software build, wrangling data, administering a system — where success is checked by whether the end result works, not by opinion.
- Does not measure
- Anything involving a visual interface or web browsing, working with people, tasks outside software and systems work.
- Known issues
- Results depend heavily on the harness wrapped around the model (the scaffold), so the same model can post very different scores; tasks lean toward software engineering.
- Trust
- high
- Last reviewed
- 2026-08-01
The test gives a model a working command-line environment and a goal — "this project won't compile, fix it," "clean up this dataset," "get this server configured" — then walks away. Success is checked mechanically: either the build passes, the data is right, the server runs, or it doesn't. No judges, no opinions. That makes it one of the cleanest measures of whether a model can do work rather than just answer questions.
How to read a score: every leaderboard entry names both a model and the harness it ran inside, and the harness matters a lot — same model, different harness, very different score. Comparing the same harness across models tells you about the models; comparing harnesses on one model tells you about the harnesses. This site records each model's best-harness number and marks who reported it.