Abstract
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation--action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98\% on RoboTwin~2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhu, X., Xu, J., Guo, Y., Wu, X., Sun, Y., Ren, X., Sun, J., Dai, Y., & Ju, X. (2026). Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling. https://omanscience.com/en/articles/dream4act-a-shared-visual-action-interface-for-multi-embodiment-video-action-modeling
MLA 9
Zhu, Xiangyu, et al. "Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling." https://omanscience.com/en/articles/dream4act-a-shared-visual-action-interface-for-multi-embodiment-video-action-modeling.
Chicago (author–date)
Zhu, Xiangyu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, and Xiaozhu Ju. 2026. "Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling." https://omanscience.com/en/articles/dream4act-a-shared-visual-action-interface-for-multi-embodiment-video-action-modeling.
Harvard
Zhu, X., Xu, J., Guo, Y., Wu, X., Sun, Y., Ren, X., Sun, J., Dai, Y. and Ju, X. (2026) 'Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling', Available at: https://omanscience.com/en/articles/dream4act-a-shared-visual-action-interface-for-multi-embodiment-video-action-modeling.
Vancouver
Zhu X, Xu J, Guo Y, Wu X, Sun Y, Ren X, et al. Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling. https://omanscience.com/en/articles/dream4act-a-shared-visual-action-interface-for-multi-embodiment-video-action-modeling
IEEE
X. Zhu, J. Xu, Y. Guo, X. Wu, Y. Sun, X. Ren, J. Sun, Y. Dai, and X. Ju, "Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling," https://omanscience.com/en/articles/dream4act-a-shared-visual-action-interface-for-multi-embodiment-video-action-modeling.