Abstract

Interactive world simulators can provide scalable environments for robot planning, policy training, and evaluation by predicting action consequences while reducing reliance on repeated physical rollouts. To serve these applications, they must generate future image sequences that respond faithfully to robot actions and preserve the dynamics of robot-object interactions over long horizons. However, existing world models typically predict the entire next latent state and often fail to capture subtle changes induced by robot actions. Such omissions can produce physically implausible outcomes, including object interpenetration and excessive deformation. To address this limitation, we propose DeltaWorld, a physically consistent interactive world simulator for robotic manipulation. Our method introduces the Delta Latent Transition Model (Delta-LTM), which predicts action-induced latent feature changes and adds them to the current latent state to obtain the next state, rather than predicting the next latent state directly. To mitigate object interpenetration and excessive deformation in predicted future frames, Interaction-aware Latent Alignment is introduced to construct counterfactual interaction regions and supervise interaction-related latent changes. DeltaWorld is evaluated on the IWS manipulation benchmark and a self-collected cross-robot dataset covering multiple robot embodiments and manipulation tasks. On the cross-robot dataset, DeltaWorld reduces FVD by 46.6% and LPIPS by 31.1% relative to the IWS baseline. These results highlight the potential of DeltaWorld for long-horizon action-conditioned video prediction in robotic manipulation.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hou, B., Cao, X., Zhang, C., Wang, S., & Cui, S. (2026). DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning. https://omanscience.com/en/articles/deltaworld-physically-consistent-interactive-world-simulators-via-action-conditioned-latent-increment-learning

MLA 9

Hou, Boyuan, et al. "DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning." https://omanscience.com/en/articles/deltaworld-physically-consistent-interactive-world-simulators-via-action-conditioned-latent-increment-learning.

Chicago (author–date)

Hou, Boyuan, Xiaoge Cao, Chaofan Zhang, Shuo Wang, and Shaowei Cui. 2026. "DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning." https://omanscience.com/en/articles/deltaworld-physically-consistent-interactive-world-simulators-via-action-conditioned-latent-increment-learning.

Harvard

Hou, B., Cao, X., Zhang, C., Wang, S. and Cui, S. (2026) 'DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning', Available at: https://omanscience.com/en/articles/deltaworld-physically-consistent-interactive-world-simulators-via-action-conditioned-latent-increment-learning.

Vancouver

Hou B, Cao X, Zhang C, Wang S, Cui S. DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning. https://omanscience.com/en/articles/deltaworld-physically-consistent-interactive-world-simulators-via-action-conditioned-latent-increment-learning

IEEE

B. Hou, X. Cao, C. Zhang, S. Wang, and S. Cui, "DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning," https://omanscience.com/en/articles/deltaworld-physically-consistent-interactive-world-simulators-via-action-conditioned-latent-increment-learning.