Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

Researchers introduce two methods to improve the routing in sparse mixture-of-experts large language models by aligning routing affinities with token-level error. They achieve improved accuracy on multiple benchmarks, while preserving the native sparse execution budget and aggregation policy.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.