الملخص
Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token. To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions (MoR) framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones. To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zheng, P., Zhang, C., Yan, J., Cao, S., Zhang, J., Wang, X., Zhang, J., Lee, J., Kim, T. H., Yang, Y., & Shen, H. T. (2026). MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models. https://omanscience.com/ar/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models
MLA 9
Zheng, Pengcheng, et al. "MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models." https://omanscience.com/ar/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.
شيكاغو (المؤلف–التاريخ)
Zheng, Pengcheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, and Heng Tao Shen. 2026. "MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models." https://omanscience.com/ar/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.
هارفارد
Zheng, P., Zhang, C., Yan, J., Cao, S., Zhang, J., Wang, X., Zhang, J., Lee, J., Kim, T. H., Yang, Y. and Shen, H. T. (2026) 'MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models', Available at: https://omanscience.com/ar/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.
فانكوفر
Zheng P, Zhang C, Yan J, Cao S, Zhang J, Wang X, et al. MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models. https://omanscience.com/ar/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models
IEEE
P. Zheng, C. Zhang, J. Yan, S. Cao, J. Zhang, X. Wang, J. Zhang, J. Lee, T. H. Kim, Y. Yang, and H. T. Shen, "MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models," https://omanscience.com/ar/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.