MADBench: Benchmarking the Security of Multi-Agent Debate
MADBench is a benchmark for evaluating the security of multi-agent debate (MAD) in large language models (LLMs). It assesses the effectiveness of MAD in mitigating or amplifying adversarial attacks. The benchmark evaluates six attack families across 356 source tasks and 3,958 test cases, showing that MAD may not improve LLM reasoning under attacks and can even amplify unauthorized actions.
Save an API key to vote.