Self-Play Search Distillation for Large Language Model Reasoning
A new framework, Self-Play Search Distillation (SPSD), generates high-quality synthetic data for Large Language Models (LLMs) using self-play of MuZero-like networks trained on board games. This technique improves LLM performance in reasoning tasks, such as mathematics, with minimal human annotation.
Save an API key to vote.