الملخص
Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Cai, Z., Wang, Y., Zhu, J., Zhu, F., & Hong, R. (2026). ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression. https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression
MLA 9
Cai, Zijing, et al. "ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression." https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.
شيكاغو (المؤلف–التاريخ)
Cai, Zijing, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, and Richang Hong. 2026. "ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression." https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.
هارفارد
Cai, Z., Wang, Y., Zhu, J., Zhu, F. and Hong, R. (2026) 'ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression', Available at: https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.
فانكوفر
Cai Z, Wang Y, Zhu J, Zhu F, Hong R. ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression. https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression
IEEE
Z. Cai, Y. Wang, J. Zhu, F. Zhu, and R. Hong, "ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression," https://omanscience.com/ar/articles/rescomemb-effective-and-efficient-multimodal-embedding-via-residual-homogeneity-compression.