GRPO Training Dynamics for Small Language Models
This paper presents a study on Group Relative Policy Optimization (GRPO) fine-tuning for small language models (SLMs) under a practical compute budget. The study analyzes the effect of group size on policy convergence, training stability, and downstream benchmark performance, and provides practical guidance for GRPO training for SLMs.
Save an API key to vote.