الباحثون

Hao Liu

المنشورات 8

نسخة أولية وصول مفتوح

GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA

Kun Wang, Yupeng Hu, Ruping Cao وآخرون · 2026

GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for gr …

نسخة أولية وصول مفتوح

Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

Shukai Gong, Xuanran Zhai, Yintianrun Zhang وآخرون · 2026

Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We pro …

نسخة أولية وصول مفتوح

VLALight: A Vision-Language-Action Model for Traffic Signal Control

Pan Zhang, Siqi Lai, Kemu Dong وآخرون · 2026

Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate percep …

نسخة أولية وصول مفتوح

Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling

Xinyu Li, Mononito Goswami, Hao Liu وآخرون · 2026

Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoni …

نسخة أولية وصول مفتوح

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu وآخرون · 2026

Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this …

نسخة أولية وصول مفتوح

On the Off-Policy Teacher in On-Policy Distillation

Langlin Huang, Hao Liu, Mononito Goswami وآخرون · 2026

On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are of …

نسخة أولية وصول مفتوح

Planned Test-Time Scaling with Coordinated Reasoning Paths

Xueqing Wu, Langxing Bai, Hritik Bansal وآخرون · 2026

Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address …

المؤلفون المشاركون