Abstract
Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Yang, P., Liu, X., Sun, J., & Tao, Q. (2026). Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis. https://omanscience.com/en/articles/revisiting-visual-representation-enhancement-of-vlms-via-kernel-canonical-correlation-analysis
MLA 9
Yang, Peilin, et al. "Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis." https://omanscience.com/en/articles/revisiting-visual-representation-enhancement-of-vlms-via-kernel-canonical-correlation-analysis.
Chicago (author–date)
Yang, Peilin, Xiaoyu Liu, Jian Sun, and Qinghua Tao. 2026. "Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis." https://omanscience.com/en/articles/revisiting-visual-representation-enhancement-of-vlms-via-kernel-canonical-correlation-analysis.
Harvard
Yang, P., Liu, X., Sun, J. and Tao, Q. (2026) 'Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis', Available at: https://omanscience.com/en/articles/revisiting-visual-representation-enhancement-of-vlms-via-kernel-canonical-correlation-analysis.
Vancouver
Yang P, Liu X, Sun J, Tao Q. Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis. https://omanscience.com/en/articles/revisiting-visual-representation-enhancement-of-vlms-via-kernel-canonical-correlation-analysis
IEEE
P. Yang, X. Liu, J. Sun, and Q. Tao, "Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis," https://omanscience.com/en/articles/revisiting-visual-representation-enhancement-of-vlms-via-kernel-canonical-correlation-analysis.