τ-Bench
SierraBenchmark for tool-using conversational agents that must coordinate with a simulated human user under domain policy, across airline, retail, telecom and banking-knowledge domains. Tasks, domain data and the evaluation harness ship together in the repository, with a public leaderboard at taubench.com. The current release, τ³-bench, adds voice and knowledge-retrieval evaluation on top of the τ² task set.
PyPI tau2 is an unrelated magnetic-relaxation package (gitlab.com/chilton-group/tau2), so no pypi artifact is declared; the Hub carries only third-party mirrors, so no huggingface_dataset either. One tier-level product covering the tau, tau2 and tau3 releases. Verified 2026-09-01 via the sierra-research/tau2-bench and tau-bench GitHub API records, both LICENSE bodies and the tau2-bench README.
Openness
5 medium confidence- license
- MIT(repository LICENSE body read on both repos
- access
- open(tasks and domain data ship in the repository, plain git clone, no gate)
- answers
- released(task definitions include the evaluation criteria
- datasheet
- yes(documents itself in the repository — domains README (structure, data format, available domains) plus arXiv:2506.07982
MIT over the whole repository, ungated, with repository documentation standing in for a Hub card as the category already accepts for mt-bench and livebench. Ladder walk: license_tier open_data (MIT) + documentation present → 5/open. Confidence medium because the answer-state reading (no held-out split) rests on the README and harness design rather than an explicit statement.
- https://raw.githubusercontent.com/sierra-research/tau2-bench/main/LICENSE recorded 2026-09-01
MIT License body, "Copyright (c) 2025 Sierra Research", standard text with no data carve-out
- https://raw.githubusercontent.com/sierra-research/tau-bench/main/LICENSE recorded 2026-09-01
MIT License body, "Copyright (c) 2024 Sierra", for the τ¹ repo
- https://raw.githubusercontent.com/sierra-research/tau2-bench/main/README.md recorded 2026-09-01
title "τ-Bench:A Benchmark for Tool-Agent-User Interaction in Real-World Domains"; τ³-bench v1.0.0/v1.0.1 announcement; arXiv:2506.07982 badge; taubench.com leaderboard; git-clone install with data in-repo; domains README link documenting domain structure and data format; backward-compatibility note naming the base split as the original τ-bench task set
Adoption
2 medium confidenceNo countable download channel exists — the only PyPI name is squatted by an unrelated package and the Hub carries no first-party dataset — so this falls back to GitHub stars, capped at level 3 by the rubric. 1,923 stars on sierra-research/tau2-bench and 1,414 on sierra-research/tau-bench, both in the 1K-10K stars band, level 2 on the stars scale. τ-bench is cited as an agent benchmark in frontier model releases, which a reviewer could read as reported_traction level 3; the measurable route is recorded instead and the question is flagged for the writer.
- https://api.github.com/repos/sierra-research/tau2-bench recorded 2026-09-01
stargazers_count 1923; license spdx MIT; homepage taubench.com; pushed 2026-09-01
- https://api.github.com/repos/sierra-research/tau-bench recorded 2026-09-01
stargazers_count 1414; license spdx MIT; description "Code and Data for Tau-Bench"
Capability
not assessedA dataset is not 'capable', so this axis is left unscored, mirroring the category pattern (gsm8k). The task suite's difficulty is a quality property, not a capability score.
- https://raw.githubusercontent.com/sierra-research/tau2-bench/main/README.md recorded 2026-09-01
a task suite plus harness; no performance, throughput or feature claim for the capability axis to read
Verified 2026-09-01