Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
A paper proposes GRAFT, an off-policy-aware framework that allows reinforcement learning models to learn from each other's experiences, improving performance and reducing the need for costly rollouts.
Save an API key to vote.