الباحثون

Henghui Ding

المنشورات 2

نسخة أولية وصول مفتوح

Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes

Runwei Guan, Rongsheng Hu, Shangshu Chen وآخرون · 2026

Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and b …

نسخة أولية وصول مفتوح

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

Qi Lyu, Jiahua Dong, Hao Shen وآخرون · 2026

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redu …

المؤلفون المشاركون