Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis

Researchers propose a new method for asynchronous reinforcement learning in large language models, addressing the issue of stale rollouts generated by earlier policies. Their method, GMC-GRPO, provides improved convergence guarantees and better performance in experiments.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.