Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training
A new method, PACE, is proposed for controlling staleness in asynchronous reinforcement learning (RL) for large language model post-training. PACE improves validation accuracy and reduces GPU time, matching synchronous RL performance.
Save an API key to vote.