الباحثون

Bohan Zhou

المنشورات 6

نسخة أولية وصول مفتوح

RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

Bohan Zhou, Xingbei Chen, Emily Huang وآخرون · 2026

Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsib …

نسخة أولية وصول مفتوح

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen وآخرون · 2026

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati …

نسخة أولية وصول مفتوح

In-Context Learning for Robots: Methods and Applications

Haojian Huang, Zexi Li, Junhao Guo وآخرون · 2026

General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize …

نسخة أولية وصول مفتوح

RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

Zexi Li, Yehang Zhang, Haojian Huang وآخرون · 2026

General-purpose robot agents must learn from experience, transfer to new tasks, and act efficiently. Code as Policies (CaP) methods generate and repair programs at runtime, incurring latency and entangling reusable mechanisms with task-specific decisions. We introduce RACaP, an agentic framework that moves coding to ev …

نسخة أولية وصول مفتوح

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

Yehang Zhang, Haojian Huang, Yifan Chang وآخرون · 2026

General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-a …

المؤلفون المشاركون