الملخص

In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence. Project page: https://GeWu-Lab.github.io/Gestalt.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yang, Z., Miao, Y., Ni, H., Chen, Z., Huang, C., Zhou, D., Chen, K., Zhang, Q., Wen, J. R., Wei, Y., & Hu, D. (2026). Gestalt: Large Multimodal Interplay Model. https://omanscience.com/ar/articles/gestalt-large-multimodal-interplay-model

MLA 9

Yang, Zequn, et al. "Gestalt: Large Multimodal Interplay Model." https://omanscience.com/ar/articles/gestalt-large-multimodal-interplay-model.

شيكاغو (المؤلف–التاريخ)

Yang, Zequn, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, Ji-Rong Wen, Yake Wei, and Di Hu. 2026. "Gestalt: Large Multimodal Interplay Model." https://omanscience.com/ar/articles/gestalt-large-multimodal-interplay-model.

هارفارد

Yang, Z., Miao, Y., Ni, H., Chen, Z., Huang, C., Zhou, D., Chen, K., Zhang, Q., Wen, J. R., Wei, Y. and Hu, D. (2026) 'Gestalt: Large Multimodal Interplay Model', Available at: https://omanscience.com/ar/articles/gestalt-large-multimodal-interplay-model.

فانكوفر

Yang Z, Miao Y, Ni H, Chen Z, Huang C, Zhou D, et al. Gestalt: Large Multimodal Interplay Model. https://omanscience.com/ar/articles/gestalt-large-multimodal-interplay-model

IEEE

Z. Yang, Y. Miao, H. Ni, Z. Chen, C. Huang, D. Zhou, K. Chen, Q. Zhang, J. R. Wen, Y. Wei, and D. Hu, "Gestalt: Large Multimodal Interplay Model," https://omanscience.com/ar/articles/gestalt-large-multimodal-interplay-model.