الملخص

Large Vision-Language Models achieve strong VQA performance, but processing high-resolution, information-rich images requires substantial computation, motivating visual token reduction. However, existing methods often prune individual tokens or rely on fixed-size cropping, limiting their ability to preserve spatially structured information such as horizontally or vertically elongated text. To address this limitation, we propose ReFIT, an instruction-guided visual token reduction framework for efficient LVLM inference. ReFIT consists of Relevance-Guided Window Reshaping (RWR) and Instruction-Guided Token Refinement (ITR), where RWR captures instruction-relevant regions by adapting to their spatial characteristics, while ITR further removes unnecessary visual tokens. Experiments on four VQA benchmarks demonstrate that ReFIT improves answer accuracy while reducing computational cost, and qualitative results demonstrate its effectiveness in localizing relevant regions and removing unnecessary visual information.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Jeong, S., Yun, J. P., & Lee, S. J. (2026). Adaptive Visual Token Reduction for Accelerated Image Understanding. https://omanscience.com/ar/articles/adaptive-visual-token-reduction-for-accelerated-image-understanding

MLA 9

Jeong, Seyoung, et al. "Adaptive Visual Token Reduction for Accelerated Image Understanding." https://omanscience.com/ar/articles/adaptive-visual-token-reduction-for-accelerated-image-understanding.

شيكاغو (المؤلف–التاريخ)

Jeong, Seyoung, Jong Pil Yun, and Sang Jun Lee. 2026. "Adaptive Visual Token Reduction for Accelerated Image Understanding." https://omanscience.com/ar/articles/adaptive-visual-token-reduction-for-accelerated-image-understanding.

هارفارد

Jeong, S., Yun, J. P. and Lee, S. J. (2026) 'Adaptive Visual Token Reduction for Accelerated Image Understanding', Available at: https://omanscience.com/ar/articles/adaptive-visual-token-reduction-for-accelerated-image-understanding.

فانكوفر

Jeong S, Yun JP, Lee SJ. Adaptive Visual Token Reduction for Accelerated Image Understanding. https://omanscience.com/ar/articles/adaptive-visual-token-reduction-for-accelerated-image-understanding

IEEE

S. Jeong, J. P. Yun, and S. J. Lee, "Adaptive Visual Token Reduction for Accelerated Image Understanding," https://omanscience.com/ar/articles/adaptive-visual-token-reduction-for-accelerated-image-understanding.