Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token. To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions (MoR) framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones. To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zheng, P., Zhang, C., Yan, J., Cao, S., Zhang, J., Wang, X., Zhang, J., Lee, J., Kim, T. H., Yang, Y., & Shen, H. T. (2026). MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models. https://omanscience.com/en/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models
MLA 9
Zheng, Pengcheng, et al. "MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models." https://omanscience.com/en/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.
Chicago (author–date)
Zheng, Pengcheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, and Heng Tao Shen. 2026. "MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models." https://omanscience.com/en/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.
Harvard
Zheng, P., Zhang, C., Yan, J., Cao, S., Zhang, J., Wang, X., Zhang, J., Lee, J., Kim, T. H., Yang, Y. and Shen, H. T. (2026) 'MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models', Available at: https://omanscience.com/en/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.
Vancouver
Zheng P, Zhang C, Yan J, Cao S, Zhang J, Wang X, et al. MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models. https://omanscience.com/en/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models
IEEE
P. Zheng, C. Zhang, J. Yan, S. Cao, J. Zhang, X. Wang, J. Zhang, J. Lee, T. H. Kim, Y. Yang, and H. T. Shen, "MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models," https://omanscience.com/en/articles/mor-mllm-mixture-of-recursions-for-efficient-multimodal-large-language-models.