الملخص
Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbf{SSP}), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Cai, Y., Rostami, M., & Thomason, J. (2026). Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models. https://omanscience.com/ar/articles/skeleton-and-strategy-prompting-training-free-negation-understanding-for-vision-language-models
MLA 9
Cai, Yuliang, et al. "Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models." https://omanscience.com/ar/articles/skeleton-and-strategy-prompting-training-free-negation-understanding-for-vision-language-models.
شيكاغو (المؤلف–التاريخ)
Cai, Yuliang, Mohammad Rostami, and Jesse Thomason. 2026. "Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models." https://omanscience.com/ar/articles/skeleton-and-strategy-prompting-training-free-negation-understanding-for-vision-language-models.
هارفارد
Cai, Y., Rostami, M. and Thomason, J. (2026) 'Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models', Available at: https://omanscience.com/ar/articles/skeleton-and-strategy-prompting-training-free-negation-understanding-for-vision-language-models.
فانكوفر
Cai Y, Rostami M, Thomason J. Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models. https://omanscience.com/ar/articles/skeleton-and-strategy-prompting-training-free-negation-understanding-for-vision-language-models
IEEE
Y. Cai, M. Rostami, and J. Thomason, "Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models," https://omanscience.com/ar/articles/skeleton-and-strategy-prompting-training-free-negation-understanding-for-vision-language-models.