Preprint Open access
MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value project …