MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

MoLE is a new framework for latent visual reasoning in vision-language models, which encourages different latent tokens to extract complementary visual information. It outperforms existing methods on five visual reasoning benchmarks, achieving an average score of 78.6.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.