Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
Researchers evaluated the effectiveness of gradient-based jailbreak detection methods in multi-turn dialogue settings. They found that existing methods can detect jailbreaks in synthetic conversations but struggle with realistic conversations. The results suggest that reliable deployment requires calibration on realistic benign conversations and consideration of various attack types and model architectures.
Save an API key to vote.