Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
Researchers propose a benchmark (NARCBench) and probing techniques for detecting collusion between AI agents in multi-agent systems, achieving high detection rates in various scenarios.
Save an API key to vote.