Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
Researchers introduce Trajectory Soup, a method to distribute mid-training budgets over multiple independent branches of large language models, improving aggregate downstream performance and extending the compute-scaling frontier of mid-training.
Save an API key to vote.