AI Potluck
Model components / Benchmark / eval datasets

τ-Bench

Sierra

Benchmark for tool-using conversational agents that must coordinate with a simulated human user under domain policy, across airline, retail, telecom and banking-knowledge domains. Tasks, domain data and the evaluation harness ship together in the repository, with a public leaderboard at taubench.com. The current release, τ³-bench, adds voice and knowledge-retrieval evaluation on top of the τ² task set.

PyPI tau2 is an unrelated magnetic-relaxation package (gitlab.com/chilton-group/tau2), so no pypi artifact is declared; the Hub carries only third-party mirrors, so no huggingface_dataset either. One tier-level product covering the tau, tau2 and tau3 releases. Verified 2026-09-01 via the sierra-research/tau2-bench and tau-bench GitHub API records, both LICENSE bodies and the tau2-bench README.

Openness

5 medium confidence
5.0
license
MIT(repository LICENSE body read on both repos
access
open(tasks and domain data ship in the repository, plain git clone, no gate)
answers
released(task definitions include the evaluation criteria
datasheet
yes(documents itself in the repository — domains README (structure, data format, available domains) plus arXiv:2506.07982

MIT over the whole repository, ungated, with repository documentation standing in for a Hub card as the category already accepts for mt-bench and livebench. Ladder walk: license_tier open_data (MIT) + documentation present → 5/open. Confidence medium because the answer-state reading (no held-out split) rests on the README and harness design rather than an explicit statement.

Adoption

2 medium confidence
2.0

No countable download channel exists — the only PyPI name is squatted by an unrelated package and the Hub carries no first-party dataset — so this falls back to GitHub stars, capped at level 3 by the rubric. 1,923 stars on sierra-research/tau2-bench and 1,414 on sierra-research/tau-bench, both in the 1K-10K stars band, level 2 on the stars scale. τ-bench is cited as an agent benchmark in frontier model releases, which a reviewer could read as reported_traction level 3; the measurable route is recorded instead and the question is flagged for the writer.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored, mirroring the category pattern (gsm8k). The task suite's difficulty is a quality property, not a capability score.

Verified 2026-09-01