Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

This paper audits tool-using AI agent benchmarks, finding defects in 34 audited tools and 12 additional tools in the AgentDojo suite. The audit treats tool interfaces as executable contracts and traces the origin of score provenance to identify defects.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.