Abstract

Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down: the top level optimizes progress toward the goal, and each level's predictions become subgoals for the planner below it. When factors in the data evolve at separated timescales, higher levels discard fast, unpredictable detail and retain slower task-relevant state. Across four simulated navigation and manipulation environments, hierarchical planning improves over a flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% using less planner compute. Ablations attribute these gains to both temporal decomposition and higher-level goal representations. With inverse-dynamics supervision, the approach extends to diverse real-robot videos from DROID, where hierarchy improves offline planning fidelity at lower planner compute.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, W., Terver, B., Rabbat, M., LeCun, Y., & Balestriero, R. (2026). H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning. https://omanscience.com/en/articles/h-jepa-end-to-end-learning-of-hierarchical-world-models-for-visual-planning

MLA 9

Zhang, Wancong, et al. "H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning." https://omanscience.com/en/articles/h-jepa-end-to-end-learning-of-hierarchical-world-models-for-visual-planning.

Chicago (author–date)

Zhang, Wancong, Basile Terver, Michael Rabbat, Yann LeCun, and Randall Balestriero. 2026. "H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning." https://omanscience.com/en/articles/h-jepa-end-to-end-learning-of-hierarchical-world-models-for-visual-planning.

Harvard

Zhang, W., Terver, B., Rabbat, M., LeCun, Y. and Balestriero, R. (2026) 'H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning', Available at: https://omanscience.com/en/articles/h-jepa-end-to-end-learning-of-hierarchical-world-models-for-visual-planning.

Vancouver

Zhang W, Terver B, Rabbat M, LeCun Y, Balestriero R. H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning. https://omanscience.com/en/articles/h-jepa-end-to-end-learning-of-hierarchical-world-models-for-visual-planning

IEEE

W. Zhang, B. Terver, M. Rabbat, Y. LeCun, and R. Balestriero, "H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning," https://omanscience.com/en/articles/h-jepa-end-to-end-learning-of-hierarchical-world-models-for-visual-planning.