الباحثون

Mingze Sun

المنشورات 2

نسخة أولية وصول مفتوح

Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

Zhenyu Liu, Zhangquan Chen, Keyi Chen وآخرون · 2026

Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-sp …

نسخة أولية وصول مفتوح

HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

Zhangquan Chen, Yaoxin Niu, Xiang An وآخرون · 2026

Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final resul …

المؤلفون المشاركون