نسخة أولية وصول مفتوح
OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing
Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representations should be expose …