Abstract
Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting plans actually contribute to the generation of action. We study one such system, a waypoint hierarchy pipeline adapted from $π_{0.5}$, and find that neither requirement is satisfied. This baseline relies on token-level autoregressive decoding (Token-AR) to generate a waypoint plan, requiring 57 very expensive vision-language model (VLM) forward passes. However, we find that erasing the waypoint endpoints has little effect on task success. Two findings reveal the misalignment of planner-executor: the planner generates outputs at an excessively fine granularity, and the executor underuses plans as a control condition. We address the latency issue with waypoint-aligned block-autoregressive decoding (Block-AR), and plan underuse issue with normalized goal modulation (NGM), a layer-wise goal path constrained by phase gating and anti-shortcut training so that the waypoint influences action generation maintaining other signals. Our method reduces the maximum number of VLM forward passes from 57 to 8 on LIBERO, including one prefix prefill, and achieves an $8.7\times$ reduction in planning latency on a Rokae dual-arm robot. With normalized goal modulation and anti-shortcut training, Block-AR's success rate on LIBERO-Long increases from 91.0% to 96.2%, while its average success rate across the four suites increases from 95.85% to 98.45%. On three bimanual tasks with this robot, success rates remain comparable across methods.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Xie, C., Ma, B., Li, G., Liu, Y., Chen, H., Zhou, X., & Yang, J. (2026). Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models. https://omanscience.com/en/articles/fast-plans-faithful-actions-closing-the-planning-execution-gap-in-hierarchical-vision-language-action-models
MLA 9
Xie, Chuanliang, et al. "Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models." https://omanscience.com/en/articles/fast-plans-faithful-actions-closing-the-planning-execution-gap-in-hierarchical-vision-language-action-models.
Chicago (author–date)
Xie, Chuanliang, Boyu Ma, Gen Li, Yizhou Liu, Houwang Chen, Xinyu Zhou, and Jianfei Yang. 2026. "Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models." https://omanscience.com/en/articles/fast-plans-faithful-actions-closing-the-planning-execution-gap-in-hierarchical-vision-language-action-models.
Harvard
Xie, C., Ma, B., Li, G., Liu, Y., Chen, H., Zhou, X. and Yang, J. (2026) 'Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models', Available at: https://omanscience.com/en/articles/fast-plans-faithful-actions-closing-the-planning-execution-gap-in-hierarchical-vision-language-action-models.
Vancouver
Xie C, Ma B, Li G, Liu Y, Chen H, Zhou X, et al. Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models. https://omanscience.com/en/articles/fast-plans-faithful-actions-closing-the-planning-execution-gap-in-hierarchical-vision-language-action-models
IEEE
C. Xie, B. Ma, G. Li, Y. Liu, H. Chen, X. Zhou, and J. Yang, "Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models," https://omanscience.com/en/articles/fast-plans-faithful-actions-closing-the-planning-execution-gap-in-hierarchical-vision-language-action-models.