Abstract

Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (\textit{e.g.}, audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer. Project page with code is available at \url{https://seohoiki3215.github.io/DCAT_project_page}

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Seo, H., Lee, B. H., Kim, M., Mah, D., Lee, J., & Chun, S. Y. (2026). Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs. https://omanscience.com/en/articles/strong-helps-weak-directional-cross-modal-alignment-transfer-in-multi-modal-llms

MLA 9

Seo, Hoigi, et al. "Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs." https://omanscience.com/en/articles/strong-helps-weak-directional-cross-modal-alignment-transfer-in-multi-modal-llms.

Chicago (author–date)

Seo, Hoigi, Byung Hyun Lee, Minjun Kim, Dohyun Mah, Jongho Lee, and Se Young Chun. 2026. "Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs." https://omanscience.com/en/articles/strong-helps-weak-directional-cross-modal-alignment-transfer-in-multi-modal-llms.

Harvard

Seo, H., Lee, B. H., Kim, M., Mah, D., Lee, J. and Chun, S. Y. (2026) 'Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs', Available at: https://omanscience.com/en/articles/strong-helps-weak-directional-cross-modal-alignment-transfer-in-multi-modal-llms.

Vancouver

Seo H, Lee BH, Kim M, Mah D, Lee J, Chun SY. Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs. https://omanscience.com/en/articles/strong-helps-weak-directional-cross-modal-alignment-transfer-in-multi-modal-llms

IEEE

H. Seo, B. H. Lee, M. Kim, D. Mah, J. Lee, and S. Y. Chun, "Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs," https://omanscience.com/en/articles/strong-helps-weak-directional-cross-modal-alignment-transfer-in-multi-modal-llms.