Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
A new technique, Bootstrapped On-Policy Self-Distillation (B-OPSD), is proposed to improve large language models. B-OPSD uses the model's own optimization progress to create a stronger self-teacher, resulting in more reliable supervision. Experiments show consistent improvements over standard OPSD on mathematical reasoning tasks.
Save an API key to vote.