Abstract

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Feng, Y., Yang, Q., Chen, R., Song, S., He, H., Zhu, M., Liu, Z., Chu, Y., Cheng, X., Wang, Y., Xu, J., & Xie, X. (2026). VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs. https://omanscience.com/en/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms

MLA 9

Feng, Yuan, et al. "VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs." https://omanscience.com/en/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.

Chicago (author–date)

Feng, Yuan, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, and Xike Xie. 2026. "VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs." https://omanscience.com/en/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.

Harvard

Feng, Y., Yang, Q., Chen, R., Song, S., He, H., Zhu, M., Liu, Z., Chu, Y., Cheng, X., Wang, Y., Xu, J. and Xie, X. (2026) 'VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs', Available at: https://omanscience.com/en/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.

Vancouver

Feng Y, Yang Q, Chen R, Song S, He H, Zhu M, et al. VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs. https://omanscience.com/en/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms

IEEE

Y. Feng, Q. Yang, R. Chen, S. Song, H. He, M. Zhu, Z. Liu, Y. Chu, X. Cheng, Y. Wang, J. Xu, and X. Xie, "VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs," https://omanscience.com/en/articles/visionweave-weaving-elastic-visual-representations-as-a-native-capability-of-mllms.