Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
Researchers propose a post-training method for multi-token prediction heads in language models, achieving similar speedup to joint pre-training with significantly fewer tokens. They also introduce a relaxation of draft token verification and an adaptive controller for dynamic MTP head engagement.
Save an API key to vote.