Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
The paper presents LOOM, a method for scaling looped mixture-of-experts (MoE) large language models (LLMs). LOOM stabilizes recurrence and diversifies computation across loops, enabling stable scaling to 9-12 loops. Experiments show improved performance in perplexity and zero-shot accuracy.
Save an API key to vote.