AI Potluck
Model components / Benchmark / eval datasets

Terminal-Bench

Laude Institute

Benchmark of difficult, hand-vetted tasks that agents must complete in a real terminal, from compiling and debugging code to system administration. It runs as a continuous benchmark with tagged releases (2.0, 2.1, 3.0) mirrored on Hugging Face and executed through the Harbor framework, with an audited leaderboard at tbench.ai. Frontier agent builders use it to track progress on long-horizon terminal work.

The org renamed from laude-institute to harbor-framework and the original repo is now terminal-bench-1; the product's numbers are the official harborframework Hub mirrors, whose 2.0 card states it is a read-only mirror of the GitHub source. PyPI terminal-bench is the legacy harness, superseded by harbor, and is left undeclared. Verified 2026-09-01 via the harbor-framework GitHub org listing, the canonical terminal-bench README, the repository LICENSE body and the harborframework Hub mirror cards.

Openness

5 high confidence
5.0
license
Apache-2.0(repository LICENSE body read (terminal-bench-1)
every version repo and all three Hub mirrors carry license
apache-2.0)
access
open(ungated Hub mirrors ("gated":false) and public GitHub task sources)
answers
released(tasks include oracle solutions and tests, run and graded locally by Harbor
datasheet
present(Hub dataset cards on each release mirror plus tbench.ai docs and repository CONTRIBUTING/task rubric)

Apache-2.0 over tasks and harness, ungated release mirrors on the Hub with cards, and oracle solutions shipped in-repo. Ladder walk: license_tier open_data (Apache-2.0) + documentation present → 5/open. Governing release recorded as the newest published mirror (3.0); all tagged releases carry the same license, so the choice does not move the score.

Adoption

3 high confidence
3.0

65,672 Hugging Face downloads in the trailing 30 days summed across the three official harborframework release mirrors (terminal-bench-2.0 37,797 + 2.1 14,616 + 3.0 13,259), which is the 10K-100K band, level 3 on the dataset scale. Third-party mirrors (zai-org/terminal-bench-2-verified at 16,739) are not counted. The legacy PyPI harness at 81,412 downloads/month is deliberately not banded — it measures harness installs, not dataset usage, and the current harness is harbor.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored, mirroring the category pattern (gsm8k).

Verified 2026-09-01