Abstract

Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency. Static consistency requires coherent 3D structure in static environments across viewpoints, while dynamic consistency requires plausible subject motion and consistent appearance over time. Geometry-aware post-training offers a promising way to improve world consistency. However, existing methods often rely on a static-scene assumption. Even those that accommodate dynamic scenes struggle to provide reliable static-consistency feedback, while dynamic consistency is often overlooked or inadequately assessed. To address these limitations, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and provides feedback by aligning each with a world prior suited to its assumptions. For static regions, WorldAlign aligns static geometry with a geometric world prior through semantically guided masked reprojection, enabling more reliable static-consistency evaluation; an auxiliary camera-motion reward discourages nearly static solutions. For dynamic subjects, WorldAlign uses a strong vision-language model (VLM) as a dynamic world prior and constructs a VLM-as-a-judge reward based on sample-specific checklists that assess dynamicity, physical plausibility, shape, and texture consistency. This decoupled design enables more effective online post-training without requiring human preference annotations. Across two pretrained image-to-video generators, Wan2.1 and Wan2.2, WorldAlign jointly improves static and dynamic consistency over existing methods without suppressing overall or subject motion. These results support decoupled world-prior alignment for more faithful visual world simulation. Project page: https://worldalign.github.io/.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

He, J., Ding, K., Tian, X., Shen, G., Ge, W., Tao, X., Wan, P., & Chen, Y. C. (2026). WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation. https://omanscience.com/en/articles/worldalign-decoupled-4d-reward-for-world-consistent-video-generation

MLA 9

He, Jing, et al. "WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation." https://omanscience.com/en/articles/worldalign-decoupled-4d-reward-for-world-consistent-video-generation.

Chicago (author–date)

He, Jing, Kaixin Ding, Xingye Tian, Guibao Shen, Wenhang Ge, Xin Tao, Pengfei Wan, and Ying-Cong Chen. 2026. "WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation." https://omanscience.com/en/articles/worldalign-decoupled-4d-reward-for-world-consistent-video-generation.

Harvard

He, J., Ding, K., Tian, X., Shen, G., Ge, W., Tao, X., Wan, P. and Chen, Y. C. (2026) 'WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation', Available at: https://omanscience.com/en/articles/worldalign-decoupled-4d-reward-for-world-consistent-video-generation.

Vancouver

He J, Ding K, Tian X, Shen G, Ge W, Tao X, et al. WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation. https://omanscience.com/en/articles/worldalign-decoupled-4d-reward-for-world-consistent-video-generation

IEEE

J. He, K. Ding, X. Tian, G. Shen, W. Ge, X. Tao, P. Wan, and Y. C. Chen, "WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation," https://omanscience.com/en/articles/worldalign-decoupled-4d-reward-for-world-consistent-video-generation.