Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
A new framework, Speculative Safety Honeypot (SSH), is proposed to proactively defend against multi-turn agent attacks. SSH uses multi-agent simulation and speculation to predict potential risks and verify them using real actions, reducing reliance on individual detection components and improving defense resilience.
Save an API key to vote.