Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
The paper introduces a metric, DSAR, to measure the safety inconsistency between reasoning traces and final answers in Large Reasoning Models (LRMs). It also proposes an RL-based method, SARA, to mitigate deceptive safety alignment in LRMs by rewarding both safe reasoning and safe final answers. Experiments show that SARA significantly reduces safety inconsistency while preserving helpfulness and utility.
Save an API key to vote.