الملخص

Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representations should be exposed to language computation for a given question? We introduce OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector. For each image-prompt pair, the projector selects a sparse set of intermediate visual depths and reuses the resulting global preference to guide both local patch fusion and dynamic visual injection into the language model. This design enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks. Across eight image-based benchmarks, OmniMoE-VL achieves an average score of 85.9 with 28B total and 9B activated parameters. Controlled comparisons show that the routed visual interface provides the dominant architectural gain, while matched route and component controls, same-image route analysis, and route interventions further support the value of coupling and question-conditioned visual access.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Qian, L., Zhu, B., Wei, J., Chen, Y., & Wang, J. (2026). OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing. https://omanscience.com/ar/articles/omnimoe-vl-a-sparse-vision-language-model-with-coupled-visual-depth-routing

MLA 9

Qian, Long, et al. "OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing." https://omanscience.com/ar/articles/omnimoe-vl-a-sparse-vision-language-model-with-coupled-visual-depth-routing.

شيكاغو (المؤلف–التاريخ)

Qian, Long, Bingke Zhu, Jiaqi Wei, Yingying Chen, and Jinqiao Wang. 2026. "OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing." https://omanscience.com/ar/articles/omnimoe-vl-a-sparse-vision-language-model-with-coupled-visual-depth-routing.

هارفارد

Qian, L., Zhu, B., Wei, J., Chen, Y. and Wang, J. (2026) 'OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing', Available at: https://omanscience.com/ar/articles/omnimoe-vl-a-sparse-vision-language-model-with-coupled-visual-depth-routing.

فانكوفر

Qian L, Zhu B, Wei J, Chen Y, Wang J. OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing. https://omanscience.com/ar/articles/omnimoe-vl-a-sparse-vision-language-model-with-coupled-visual-depth-routing

IEEE

L. Qian, B. Zhu, J. Wei, Y. Chen, and J. Wang, "OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing," https://omanscience.com/ar/articles/omnimoe-vl-a-sparse-vision-language-model-with-coupled-visual-depth-routing.