Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

Researchers propose a post-training method for multi-token prediction heads in language models, achieving similar speedup to joint pre-training with significantly fewer tokens. They also introduce a relaxation of draft token verification and an adaptive controller for dynamic MTP head engagement.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.