AI Potluck
Model components / Benchmark / eval datasets

AgentDojo

ETH Zurich SPY Lab

Benchmark for evaluating prompt-injection attacks and defenses on tool-using LLM agents. It ships realistic task suites across workspace, banking, travel and Slack environments together with injection test cases, and grades agents on both task completion and security. The suite installs from PyPI and publishes results for common models and defenses.

Boundary call: a prompt-injection benchmark filed as benchmark data rather than a safeguard, because the product is an evaluation suite. Verified 2026-09-01 via the ethz-spylab/agentdojo GitHub API record, the MIT LICENSE body, the repository README and the PyPI project JSON, whose project_urls point back at the repository.

Openness

5 high confidence
5.0
license
MIT(repository LICENSE body read
access
open(pip install agentdojo
answers
released(ground-truth task and security checks ship with the suite and grade locally
datasheet
present(documentation site agentdojo.spylab.ai with API and suite docs, published results pages, plus arXiv:2406.13352)

MIT over the whole repository, ungated distribution through PyPI and GitHub, and a maintained documentation site standing in for a Hub card under the category's repo-documentation precedent. Ladder walk: license_tier open_data (MIT) + documentation present → 5/open.

Adoption

3 medium confidence
3.0

24,142 PyPI downloads in the last month for the agentdojo package, which ships the benchmark suite itself and is backlink-verified to the repository. On the dataset scale that is the 10K-100K band, level 3. The unit caveat is flagged: the dataset bands are calibrated on HF dataset downloads and this product's channel is PyPI, but the suite IS the distribution (no HF dataset exists), matching how evalplus is banded in this category.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored, mirroring the category pattern (gsm8k).

Verified 2026-09-01