الباحثون

Yikai Wang

المنشورات 8

نسخة أولية وصول مفتوح

CADForge: Agentic Single-View CAD Reconstruction with Explicit Geometry Reasoning

Keyang Lu, Zhifei Yang, Tianao Dong وآخرون · 2026

Reconstructing editable parametric CAD models from a single-view image is of great practical value for modern manufacturing, yet remains challenging due to incomplete geometric observations and complex inter-part relationships. To address it, we propose CADForge, an agentic framework that progressively converts a singl …

نسخة أولية وصول مفتوح

Native Action-Prior Learning from Videos for World Action Models

Zhaochong An, Fei Zhang, Menglin Jia وآخرون · 2026

World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual rep …

نسخة أولية وصول مفتوح

World Action Modeling with Progressive Visual Planning

Fei Zhang, Zhaochong An, Duncan Frost وآخرون · 2026

World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WA …

نسخة أولية وصول مفتوح

KPI: A Promptable Kernel for Physical Interaction on Humanoids

Yikai Wang, Honghao Zhu, Xiao Hu وآخرون · 2026

Humanoids now walk, balance and reach with remarkable generality: one whole-body tracking policy follows references from a human, or from an end-to-end policy. That generality travels in the trajectory, and a trajectory alone carries limited information about the interaction it should produce: at contact, the executing …

نسخة أولية وصول مفتوح

In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

Yikai Wang, Xiao Han, Mengmeng Xu وآخرون · 2026

Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by addition …

نسخة أولية وصول مفتوح

Anticipatory Robot Goalkeeping via Monotone Optimal Stopping

Hao E. Zhang, Ruize Geng, Yisen Li وآخرون · 2026

Robots engaged in fast physical interactions often need to act before the intent of another agent is fully known. Anticipatory goalkeeping illustrates this challenge. Waiting provides more reliable information about the target but reduces the physical opportunity for interception, whereas acting early preserves reachab …

نسخة أولية وصول مفتوح

Dynamics-Induced Commitment in Learning-Based Robotic Penalty Kicks

Ruize Geng, Hao E. Zhang, Yisen Li وآخرون · 2026

Learning in robotic games is constrained not only by strategic information but also by what the body can still execute. We study this coupling in a hierarchical humanoid-quadruped penalty system in which game-level self-play policies command fixed soccer whole-body controllers (S-WBCs). The humanoid shooting skill is i …

المؤلفون المشاركون