Abstract

Video demonstrations offer a scalable alternative to costly robot data for learning manipulation, yet existing reconstruction-based approaches often rely on constrained camera viewpoints or human-to-robot retargeting, while the reconstructed trajectories are difficult to adapt to new objects configurations without distorting the trajectory shape. Another key limitation is that the resulting policies often lack precise object-level 3D geometry awareness, limiting object grounding and object shape awareness critical for precise manipulation. To bridge these gaps, we propose VidAct, an efficient video-to-robot framework that learns object-centric, 3D-aware manipulation policies from a single monocular video per task and enables zero-shot real-world deployment. VidAct consists of three key components. First, VidAct reconstructs object meshes and motion from arbitrary demo videos and canonicalizes the motion in the static object frame, avoiding embodiment-specific retargeting and accommodating diverse camera viewpoints. Second, VidAct employ residual trajectory transfer for adapting the reconstructed motion to novel object configurations while preserving its motion shape. Finally, as the key policy-learning component, VidAct predicts simulation-provided privileged complete-object point clouds at each frame as an auxiliary task while retaining RGB-only deployment, providing dense object-centric supervision over both object pose and 3D geometry. Experiments on human, robot, generated, and internet videos demonstrate broad video applicability and zero-shot deployment. Per-frame complete-object 3D supervision improves policy generalization and sim-to-real success, while residual trajectory transfer enables reliable trajectory adaptation with better shape preservation.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, H., Zhang, M., Wu, Z., Tian, Y., Chen, D., Shen, F., Meng, Y., Yao, X., Zhang, H., Liu, Z., Bing, Z., & Knoll, A. (2026). VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness. https://omanscience.com/en/articles/vidact-learning-manipulation-from-in-the-wild-videos-with-object-centric-3d-awareness

MLA 9

Li, Hang, et al. "VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness." https://omanscience.com/en/articles/vidact-learning-manipulation-from-in-the-wild-videos-with-object-centric-3d-awareness.

Chicago (author–date)

Li, Hang, Mingxin Zhang, Zihan Wu, Yang Tian, Dong Chen, Fengyi Shen, Yuan Meng, Xiangtong Yao, Heng Zhang, Ziyuan Liu, Zhenshan Bing, and Alois Knoll. 2026. "VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness." https://omanscience.com/en/articles/vidact-learning-manipulation-from-in-the-wild-videos-with-object-centric-3d-awareness.

Harvard

Li, H., Zhang, M., Wu, Z., Tian, Y., Chen, D., Shen, F., Meng, Y., Yao, X., Zhang, H., Liu, Z., Bing, Z. and Knoll, A. (2026) 'VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness', Available at: https://omanscience.com/en/articles/vidact-learning-manipulation-from-in-the-wild-videos-with-object-centric-3d-awareness.

Vancouver

Li H, Zhang M, Wu Z, Tian Y, Chen D, Shen F, et al. VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness. https://omanscience.com/en/articles/vidact-learning-manipulation-from-in-the-wild-videos-with-object-centric-3d-awareness

IEEE

H. Li, M. Zhang, Z. Wu, Y. Tian, D. Chen, F. Shen, Y. Meng, X. Yao, H. Zhang, Z. Liu, Z. Bing, and A. Knoll, "VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness," https://omanscience.com/en/articles/vidact-learning-manipulation-from-in-the-wild-videos-with-object-centric-3d-awareness.