الباحثون

Ivan Laptev

المنشورات 9

نسخة أولية وصول مفتوح

iGPC: Generative Motion Priors for Object-Aware Humanoid Interaction

Humanoid robots operating in unstructured environments must combine robust whole-body control with the ability to perceive and physically interact with surrounding objects. While large-scale human motion data provides powerful priors for natural and versatile humanoid control, effectively transferring such priors to pe …

نسخة أولية وصول مفتوح

Streaming Multi-Track Timeline Control for 3D Human Motion Generation

Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation …

نسخة أولية وصول مفتوح

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosi …

نسخة أولية وصول مفتوح

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this …

نسخة أولية وصول مفتوح

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

Qi Lyu, Jiahua Dong, Hao Shen وآخرون · 2026

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redu …

نسخة أولية وصول مفتوح

PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end mod …

نسخة أولية وصول مفتوح

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing

Sohyun Lee, Yoonjae Baek, Jaesang Won وآخرون · 2026

Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VL …

نسخة أولية وصول مفتوح

The Robot Data Factory

Sami Haddadin, Ivan Laptev, Ian Reid وآخرون · 2026

Physical AI requires more than increasingly large robot datasets: intelligent robots acquire knowledge through continuous interaction with the physical world. We argue that the defining scientific resource of Physical AI is therefore not raw robot data alone, but robot experience - physically grounded interaction whose …

نسخة أولية وصول مفتوح

World-Action Models for Robot Learning and Control: A Survey

Zuxing Lu, Hongjia Zhai, Guanzhi Wang وآخرون · 2026

Robots operating in open environments act under partial observability, physical constraints, and dynamic task contexts. Beyond mapping observations and language instructions to actions, they must anticipate how candidate actions may affect future states and task-relevant outcomes. Recent advances in world models, video …

المؤلفون المشاركون