الباحثون

Hongkun Cao

المنشورات 2

نسخة أولية وصول مفتوح

HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

Zhangquan Chen, Yaoxin Niu, Xiang An وآخرون · 2026

Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final resul …

المؤلفون المشاركون