CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

A new method called CARM (Cancellation-Aware Response Masking) is proposed to address off-policy issues in large language model (LLM) reinforcement learning. CARM improves upon existing sequence-level masking by taking the absolute value of token log-ratios, preventing policy drift and resulting in better performance on mathematical reasoning and code generation tasks.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.