Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models
Researchers introduce two methods to improve the routing in sparse mixture-of-experts large language models by aligning routing affinities with token-level error. They achieve improved accuracy on multiple benchmarks, while preserving the native sparse execution budget and aggregation policy.
Save an API key to vote.