الملخص

Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching. Mainstream approaches primarily adopt independent encoding of vision and language modality followed by rigid late fusion such as concatenation, inherently omitting native vision-language interactions and introducing cross-modal semantic distortion. To address this challenge, we propose OpticalRec, the first visual-space unified encoding paradigm for multimodal collaborative filtering, a fundamental recommendation setting. Instead of isolated modality-specific encoding, OpticalRec renders item textual metadata as visual glyphs, enabling native image-text interaction within the visual encoder - the perceptual encoding level. The resulting representations are further processed by the language decoder - the semantic encoding level, allowing OpticalRec to exploit the dual-attention mechanism of modern vision-language models that previous encoding methods omitted. OpticalRec's efficacy is theoretically supported by mutual information analysis and empirically demonstrated through superior performance across strong baselines and benchmarks. As a plug-and-play module, OpticalRec (1) introduces minimal cost, (2) is robust against rendered text font, color and layout, etc., and (3) integrates seamlessly into existing multimodal collaborative filtering models.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wang, Y., Guo, Z., Hou, Y., Wang, Y., Kim, K., Yue, Z., Xing, S., Li, H., Xia, H., Zhang, R., Tu, Z., & McAuley, J. (2026). OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation. https://omanscience.com/ar/articles/opticalrec-unified-optical-vision-language-representation-for-multimodal-recommendation

MLA 9

Wang, Yueqi, et al. "OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation." https://omanscience.com/ar/articles/opticalrec-unified-optical-vision-language-representation-for-multimodal-recommendation.

شيكاغو (المؤلف–التاريخ)

Wang, Yueqi, Zitian Guo, Yupeng Hou, Yifei Wang, Kibum Kim, Zhenrui Yue, Shuo Xing, Haodong Li, Heming Xia, Renrui Zhang, Zhengzhong Tu, and Julian McAuley. 2026. "OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation." https://omanscience.com/ar/articles/opticalrec-unified-optical-vision-language-representation-for-multimodal-recommendation.

هارفارد

Wang, Y., Guo, Z., Hou, Y., Wang, Y., Kim, K., Yue, Z., Xing, S., Li, H., Xia, H., Zhang, R., Tu, Z. and McAuley, J. (2026) 'OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation', Available at: https://omanscience.com/ar/articles/opticalrec-unified-optical-vision-language-representation-for-multimodal-recommendation.

فانكوفر

Wang Y, Guo Z, Hou Y, Wang Y, Kim K, Yue Z, et al. OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation. https://omanscience.com/ar/articles/opticalrec-unified-optical-vision-language-representation-for-multimodal-recommendation

IEEE

Y. Wang, Z. Guo, Y. Hou, Y. Wang, K. Kim, Z. Yue, S. Xing, H. Li, H. Xia, R. Zhang, Z. Tu, and J. McAuley, "OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation," https://omanscience.com/ar/articles/opticalrec-unified-optical-vision-language-representation-for-multimodal-recommendation.