Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
Researchers developed DART, a runtime framework that detects and attributes representation shifts in multi-turn LLM agents, reducing attack success rates and outperforming existing defenses. The framework uses denoising to identify and intervene in potentially harmful behavior.
Save an API key to vote.