الباحثون

Zhiyuan Gao

المنشورات 5

نسخة أولية وصول مفتوح

Reliability-Aware Future Conditioning for Temporally Robust Robot Manipulation

A generated video of a task the robot is about to perform is useful guidance only if it depicts the phase the robot is actually in. We show that temporal misalignment can turn a task-consistent generated future into actively harmful guidance. On CALVIN, a five-frame early shift nearly erases the benefit of generated fu …

نسخة أولية وصول مفتوح

CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding

Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed. Existing fault-recovery methods often require task-specific retraining, fault labels, explicit diagnosis, or privileged embodiment information. We …

نسخة أولية وصول مفتوح

HEIR: Learning Human-Entity Interactions with Functional Roles

Di Wen, Wenhao Guo, Yuedong Tan وآخرون · 2026

Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score indiv …

نسخة أولية وصول مفتوح

FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

Zhiyuan Gao, Di Wen, Yanxiang Zhan وآخرون · 2026

Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in …

نسخة أولية وصول مفتوح

KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

Learning robot manipulation policies typically requires substantial demonstration data, which are costly to collect on real robots. Recent methods generate robot demonstrations from human videos by adapting recovered motion and validating the resulting trajectories in simulation. However, methods centered on motion-ref …

المؤلفون المشاركون