MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems
MAADBench is a new benchmark for anomaly detection in multi-agent systems (MAS) with evolving LLM backbones. It addresses challenges in MAS AD benchmarking by providing a refreshable paradigm with sampled-and-coupled generative tasks, refreshable trace generation, and automated label provision. The benchmark is designed to support diverse LLM backbones and offers a rich research agenda for MAS-specific anomaly detection.
Save an API key to vote.