Abstract
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Dai, Q., Wen, L., Duan, J., Dai, Y., Wang, D., Wang, M., Wang, M., Liu, J., Yan, H., & Kang, Z. (2026). HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling. https://omanscience.com/en/articles/hril-learning-multimodal-synergy-via-higher-order-tensor-modeling
MLA 9
Dai, Qun, et al. "HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling." https://omanscience.com/en/articles/hril-learning-multimodal-synergy-via-higher-order-tensor-modeling.
Chicago (author–date)
Dai, Qun, Liangjian Wen, Jiang Duan, Yong Dai, Dongkai Wang, Maolin Wang, Mingjie Wang, Jianzhuang Liu, He Yan, and Zhao Kang. 2026. "HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling." https://omanscience.com/en/articles/hril-learning-multimodal-synergy-via-higher-order-tensor-modeling.
Harvard
Dai, Q., Wen, L., Duan, J., Dai, Y., Wang, D., Wang, M., Wang, M., Liu, J., Yan, H. and Kang, Z. (2026) 'HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling', Available at: https://omanscience.com/en/articles/hril-learning-multimodal-synergy-via-higher-order-tensor-modeling.
Vancouver
Dai Q, Wen L, Duan J, Dai Y, Wang D, Wang M, et al. HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling. https://omanscience.com/en/articles/hril-learning-multimodal-synergy-via-higher-order-tensor-modeling
IEEE
Q. Dai, L. Wen, J. Duan, Y. Dai, D. Wang, M. Wang, M. Wang, J. Liu, H. Yan, and Z. Kang, "HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling," https://omanscience.com/en/articles/hril-learning-multimodal-synergy-via-higher-order-tensor-modeling.