Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

Researchers propose a benchmark (NARCBench) and probing techniques for detecting collusion between AI agents in multi-agent systems, achieving high detection rates in various scenarios.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.