Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue

Researchers evaluated the effectiveness of gradient-based jailbreak detection methods in multi-turn dialogue settings. They found that existing methods can detect jailbreaks in synthetic conversations but struggle with realistic conversations. The results suggest that reliable deployment requires calibration on realistic benign conversations and consideration of various attack types and model architectures.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.