Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
This paper audits tool-using AI agent benchmarks, finding defects in 34 audited tools and 12 additional tools in the AgentDojo suite. The audit treats tool interfaces as executable contracts and traces the origin of score provenance to identify defects.
Save an API key to vote.