[
    {
        "id": "osp-19621",
        "type": "article-journal",
        "title": "Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis",
        "author": [
            {
                "family": "Yang",
                "given": "Peilin"
            },
            {
                "family": "Liu",
                "given": "Xiaoyu"
            },
            {
                "family": "Sun",
                "given": "Jian"
            },
            {
                "family": "Tao",
                "given": "Qinghua"
            }
        ],
        "URL": "https://omanscience.com/en/articles/revisiting-visual-representation-enhancement-of-vlms-via-kernel-canonical-correlation-analysis",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance."
    }
]