Abstract

World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution, whereas wrist cameras move with the end effector, mixing local interaction changes with viewpoint shifts and self-occlusion. These contrasting predictive demands suggest that the two views may benefit from different future objectives. We introduce FutureDuet, which retains both views for control, while allowing each visual stream to receive a different future objective. For the main view, future RGB models task evolution, while interaction masks and robot skeletons focus supervision on task objects and robot motion. For the wrist stream, future latent prediction models short-horizon interaction changes without requiring pixel-level reconstruction. ActionDiT jointly reads the resulting Task State and Interaction State, combining scene-level progress with close-range interaction evidence. All auxiliary prediction modules are training-only, adding no inference overhead. FutureDuet achieves 94.2% clean and 94.1% randomized success on RoboTwin50 and 99.2% average success on LIBERO. The improvements are most pronounced on six RoboTwin50 tasks that require precise interaction, averaging gains of 9.2% and 12.8% over Fast-WAM in clean and randomized settings. Controlled studies further show complementary gains from separating the wrist pathway and designing future supervision separately for the two views.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wu, J., Huang, Y., Liu, J., Zhang, W., Huang, H., Chen, Y., Jiang, J., & Zhang, C. (2026). FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models. https://omanscience.com/en/articles/futureduet-decoupling-observation-access-from-future-supervision-in-world-action-models

MLA 9

Wu, Jie, et al. "FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models." https://omanscience.com/en/articles/futureduet-decoupling-observation-access-from-future-supervision-in-world-action-models.

Chicago (author–date)

Wu, Jie, Yuzhi Huang, Junqi Liu, Weichen Zhang, Haibin Huang, Yin Chen, Jingyan Jiang, and Chi Zhang. 2026. "FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models." https://omanscience.com/en/articles/futureduet-decoupling-observation-access-from-future-supervision-in-world-action-models.

Harvard

Wu, J., Huang, Y., Liu, J., Zhang, W., Huang, H., Chen, Y., Jiang, J. and Zhang, C. (2026) 'FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models', Available at: https://omanscience.com/en/articles/futureduet-decoupling-observation-access-from-future-supervision-in-world-action-models.

Vancouver

Wu J, Huang Y, Liu J, Zhang W, Huang H, Chen Y, et al. FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models. https://omanscience.com/en/articles/futureduet-decoupling-observation-access-from-future-supervision-in-world-action-models

IEEE

J. Wu, Y. Huang, J. Liu, W. Zhang, H. Huang, Y. Chen, J. Jiang, and C. Zhang, "FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models," https://omanscience.com/en/articles/futureduet-decoupling-observation-access-from-future-supervision-in-world-action-models.