نسخة أولية وصول مفتوح
Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs
Vision-language models (VLMs) have achieved strong performance on tasks such as visual question answering, yet small image resizes can turn correct answers into errors. We investigate whether changes in visual configuration, such as image tiling and token arrangement, contribute to this instability. Across seven checkp …