الملخص

Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Cai, Z., Wang, Y., Zhu, J., Zhu, F., & Hong, R. (2026). ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression. https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression

MLA 9

Cai, Zijing, et al. "ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression." https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.

شيكاغو (المؤلف–التاريخ)

Cai, Zijing, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, and Richang Hong. 2026. "ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression." https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.

هارفارد

Cai, Z., Wang, Y., Zhu, J., Zhu, F. and Hong, R. (2026) 'ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression', Available at: https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.

فانكوفر

Cai Z, Wang Y, Zhu J, Zhu F, Hong R. ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression. https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression

IEEE

Z. Cai, Y. Wang, J. Zhu, F. Zhu, and R. Hong, "ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression," https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.