نسخة أولية وصول مفتوح
Seek-and-View Reasoning for Multi-View Spatial Understanding
Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, w …