نسخة أولية وصول مفتوح
SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention …