AI Change Tracker

Terminal-Bench

https://www.tbench.ai/leaderboard

Kind
benchmark
Maintainer
Stanford / Laude Institute
Feeds axes
agentic
Measures
Whether a model can complete real multi-step technical tasks on its own — fixing a broken software build, wrangling data, administering a system — where success is checked by whether the end result works, not by opinion.
Does not measure
Anything involving a visual interface or web browsing, working with people, tasks outside software and systems work.
Known issues
Results depend heavily on the harness wrapped around the model (the scaffold), so the same model can post very different scores; tasks lean toward software engineering.
Trust
high
Last reviewed
2026-08-01

The test gives a model a working command-line environment and a goal — "this project won't compile, fix it," "clean up this dataset," "get this server configured" — then walks away. Success is checked mechanically: either the build passes, the data is right, the server runs, or it doesn't. No judges, no opinions. That makes it one of the cleanest measures of whether a model can do work rather than just answer questions.

How to read a score: every leaderboard entry names both a model and the harness it ran inside, and the harness matters a lot — same model, different harness, very different score. Comparing the same harness across models tells you about the models; comparing harnesses on one model tells you about the harnesses. This site records each model's best-harness number and marks who reported it.

← All methods