Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training

A new method, PACE, is proposed for controlling staleness in asynchronous reinforcement learning (RL) for large language model post-training. PACE improves validation accuracy and reduces GPU time, matching synchronous RL performance.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.