Abstract

Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We propose Visual Goal-conditioned Action Reasoning (ViGAR), a hierarchical framework that factorizes manipulation into a visual subgoal planner and a subgoal executor. Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal. Both components share a pretrained world-model representation, enabling task-level planning and action generation to benefit from common physical knowledge. Moreover, our framework naturally supports in-context learning: using a global goal image as context can induce different subtask decompositions and behaviors without parameter updates. On the RoboTwin Clean2Random benchmark, ViGAR achieves 82.00% and 67.02% success rates under the Clean and Random settings, respectively, surpassing the strongest baseline by 12.86 percentage points in average success rate. Real-world robot experiments on five compositional and two in-context learning tasks further confirm the effectiveness of ViGAR.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Gong, S., Zhai, X., Zhang, Y., Cui, R., Huang, Y., Fu, Y., Lyu, D., Li, C., Song, X., Lin, P., Wang, C., Jia, M., Deng, Y., Fang, J., Liang, B., Li, J., Gao, Y., Liu, H., & Zhou, D. (2026). Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation. https://omanscience.com/en/articles/rethinking-world-action-model-for-compositional-and-in-context-robotic-manipulation

MLA 9

Gong, Shukai, et al. "Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation." https://omanscience.com/en/articles/rethinking-world-action-model-for-compositional-and-in-context-robotic-manipulation.

Chicago (author–date)

Gong, Shukai, Xuanran Zhai, Yintianrun Zhang, Ruopeng Cui, Ye Huang, Yiyang Fu, Dexuan Lyu, Chaojie Li, Xinyi Song, Peiwen Lin, Chuang Wang, Mingyuan Jia, Yufan Deng, Jiaxin Fang, Bo Liang, Jiaxin Li, Yuxiang Gao, Hao Liu, and Daquan Zhou. 2026. "Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation." https://omanscience.com/en/articles/rethinking-world-action-model-for-compositional-and-in-context-robotic-manipulation.

Harvard

Gong, S., Zhai, X., Zhang, Y., Cui, R., Huang, Y., Fu, Y., Lyu, D., Li, C., Song, X., Lin, P., Wang, C., Jia, M., Deng, Y., Fang, J., Liang, B., Li, J., Gao, Y., Liu, H. and Zhou, D. (2026) 'Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation', Available at: https://omanscience.com/en/articles/rethinking-world-action-model-for-compositional-and-in-context-robotic-manipulation.

Vancouver

Gong S, Zhai X, Zhang Y, Cui R, Huang Y, Fu Y, et al. Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation. https://omanscience.com/en/articles/rethinking-world-action-model-for-compositional-and-in-context-robotic-manipulation

IEEE

S. Gong, X. Zhai, Y. Zhang, R. Cui, Y. Huang, Y. Fu, D. Lyu, C. Li, X. Song, P. Lin, C. Wang, M. Jia, Y. Deng, J. Fang, B. Liang, J. Li, Y. Gao, H. Liu, and D. Zhou, "Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation," https://omanscience.com/en/articles/rethinking-world-action-model-for-compositional-and-in-context-robotic-manipulation.