الملخص
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Feng, Y., Yang, Q., Chen, R., Song, S., He, H., Zhu, M., Liu, Z., Chu, Y., Cheng, X., Wang, Y., Xu, J., & Xie, X. (2026). VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs. https://omanscience.com/ar/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms
MLA 9
Feng, Yuan, et al. "VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs." https://omanscience.com/ar/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.
شيكاغو (المؤلف–التاريخ)
Feng, Yuan, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, and Xike Xie. 2026. "VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs." https://omanscience.com/ar/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.
هارفارد
Feng, Y., Yang, Q., Chen, R., Song, S., He, H., Zhu, M., Liu, Z., Chu, Y., Cheng, X., Wang, Y., Xu, J. and Xie, X. (2026) 'VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs', Available at: https://omanscience.com/ar/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.
فانكوفر
Feng Y, Yang Q, Chen R, Song S, He H, Zhu M, et al. VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs. https://omanscience.com/ar/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms
IEEE
Y. Feng, Q. Yang, R. Chen, S. Song, H. He, M. Zhu, Z. Liu, Y. Chu, X. Cheng, Y. Wang, J. Xu, and X. Xie, "VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs," https://omanscience.com/ar/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.