الباحثون

Drandreb Earl Juanico

المنشورات 1

نسخة أولية وصول مفتوح

Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models

VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We study this readout/recovery separation in Qwen3-VL-4B-Instruct …