الملخص
Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Niu, Y., Chen, Z., Zhang, Y., An, X., Wang, Z., Liao, C. T., Cao, H., & Huang, R. (2026). EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception. https://omanscience.com/ar/articles/evivit-evidence-adaptive-vision-transformers-for-fine-grained-perception
MLA 9
Niu, Yaoxin, et al. "EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception." https://omanscience.com/ar/articles/evivit-evidence-adaptive-vision-transformers-for-fine-grained-perception.
شيكاغو (المؤلف–التاريخ)
Niu, Yaoxin, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, and Ruqi Huang. 2026. "EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception." https://omanscience.com/ar/articles/evivit-evidence-adaptive-vision-transformers-for-fine-grained-perception.
هارفارد
Niu, Y., Chen, Z., Zhang, Y., An, X., Wang, Z., Liao, C. T., Cao, H. and Huang, R. (2026) 'EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception', Available at: https://omanscience.com/ar/articles/evivit-evidence-adaptive-vision-transformers-for-fine-grained-perception.
فانكوفر
Niu Y, Chen Z, Zhang Y, An X, Wang Z, Liao CT, et al. EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception. https://omanscience.com/ar/articles/evivit-evidence-adaptive-vision-transformers-for-fine-grained-perception
IEEE
Y. Niu, Z. Chen, Y. Zhang, X. An, Z. Wang, C. T. Liao, H. Cao, and R. Huang, "EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception," https://omanscience.com/ar/articles/evivit-evidence-adaptive-vision-transformers-for-fine-grained-perception.