نسخة أولية وصول مفتوح
Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models
Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this assumption often fails, leaving a retrieval-reading gap: evidence found but not used. A paired protocol isolates this gap by comparing answers from the retrieved page i …