نسخة أولية وصول مفتوح
GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions
Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on indi …