[
    {
        "id": "osp-25868",
        "type": "article-journal",
        "title": "Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models",
        "author": [
            {
                "family": "Xie",
                "given": "Chuanliang"
            },
            {
                "family": "Ma",
                "given": "Boyu"
            },
            {
                "family": "Li",
                "given": "Gen"
            },
            {
                "family": "Liu",
                "given": "Yizhou"
            },
            {
                "family": "Chen",
                "given": "Houwang"
            },
            {
                "family": "Zhou",
                "given": "Xinyu"
            },
            {
                "family": "Yang",
                "given": "Jianfei"
            }
        ],
        "URL": "https://omanscience.com/en/articles/fast-plans-faithful-actions-closing-the-planning-execution-gap-in-hierarchical-vision-language-action-models",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting plans actually contribute to the generation of action. We study one such system, a waypoint hierarchy pipeline adapted from $π_{0.5}$, and find that neither requirement is satisfied. This baseline relies on token-level autoregressive decoding (Token-AR) to generate a waypoint plan, requiring 57 very expensive vision-language model (VLM) forward passes. However, we find that erasing the waypoint endpoints has little effect on task success. Two findings reveal the misalignment of planner-executor: the planner generates outputs at an excessively fine granularity, and the executor underuses plans as a control condition. We address the latency issue with waypoint-aligned block-autoregressive decoding (Block-AR), and plan underuse issue with normalized goal modulation (NGM), a layer-wise goal path constrained by phase gating and anti-shortcut training so that the waypoint influences action generation maintaining other signals. Our method reduces the maximum number of VLM forward passes from 57 to 8 on LIBERO, including one prefix prefill, and achieves an $8.7\\times$ reduction in planning latency on a Rokae dual-arm robot. With normalized goal modulation and anti-shortcut training, Block-AR's success rate on LIBERO-Long increases from 91.0% to 96.2%, while its average success rate across the four suites increases from 95.85% to 98.45%. On three bimanual tasks with this robot, success rates remain comparable across methods."
    }
]