Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

The paper introduces a metric, DSAR, to measure the safety inconsistency between reasoning traces and final answers in Large Reasoning Models (LRMs). It also proposes an RL-based method, SARA, to mitigate deceptive safety alignment in LRMs by rewarding both safe reasoning and safe final answers. Experiments show that SARA significantly reduces safety inconsistency while preserving helpfulness and utility.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.