الملخص

Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Chen, K., Hou, Y., Liu, H., & Jia, Y. (2026). Unsupervised Long-Tailed Adaptation of Vision-Language Models. https://omanscience.com/ar/articles/unsupervised-long-tailed-adaptation-of-vision-language-models

MLA 9

Chen, Keliang, et al. "Unsupervised Long-Tailed Adaptation of Vision-Language Models." https://omanscience.com/ar/articles/unsupervised-long-tailed-adaptation-of-vision-language-models.

شيكاغو (المؤلف–التاريخ)

Chen, Keliang, Yaxin Hou, Hui Liu, and Yuheng Jia. 2026. "Unsupervised Long-Tailed Adaptation of Vision-Language Models." https://omanscience.com/ar/articles/unsupervised-long-tailed-adaptation-of-vision-language-models.

هارفارد

Chen, K., Hou, Y., Liu, H. and Jia, Y. (2026) 'Unsupervised Long-Tailed Adaptation of Vision-Language Models', Available at: https://omanscience.com/ar/articles/unsupervised-long-tailed-adaptation-of-vision-language-models.

فانكوفر

Chen K, Hou Y, Liu H, Jia Y. Unsupervised Long-Tailed Adaptation of Vision-Language Models. https://omanscience.com/ar/articles/unsupervised-long-tailed-adaptation-of-vision-language-models

IEEE

K. Chen, Y. Hou, H. Liu, and Y. Jia, "Unsupervised Long-Tailed Adaptation of Vision-Language Models," https://omanscience.com/ar/articles/unsupervised-long-tailed-adaptation-of-vision-language-models.