Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money
A new benchmark and detector for agentic commerce fraud in AI agents that spend money, including a taxonomy, a benchmark, and an open-source detector stack. The benchmark consists of 20 fraud classes generated from production aggregates, and the detector stack can audit an agent configuration, replay hostile counterparties, and run detectors inline.
Save an API key to vote.