CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
A new method called CARM (Cancellation-Aware Response Masking) is proposed to address off-policy issues in large language model (LLM) reinforcement learning. CARM improves upon existing sequence-level masking by taking the absolute value of token log-ratios, preventing policy drift and resulting in better performance on mathematical reasoning and code generation tasks.
Save an API key to vote.