[
    {
        "id": "osp-18166",
        "type": "article-journal",
        "title": "PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies",
        "author": [
            {
                "family": "Huang",
                "given": "Qing"
            },
            {
                "family": "Yang",
                "given": "Yifei"
            },
            {
                "family": "Zou",
                "given": "Ziqing"
            },
            {
                "family": "Chen",
                "given": "Anzhe"
            },
            {
                "family": "Zhu",
                "given": "Zhenjie"
            },
            {
                "family": "Wei",
                "given": "Yufei"
            },
            {
                "family": "Xiong",
                "given": "Rong"
            },
            {
                "family": "Wang",
                "given": "Yue"
            }
        ],
        "URL": "https://omanscience.com/en/articles/pathtime-vla-path-time-decoupling-for-factorized-post-training-of-vision-language-action-policies",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace. This coupling complicates adaptation from teleoperation: useful geometric guidance comes with timing shaped by interface delays and operator behavior. Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA. We introduce PathTime-VLA, which represents motion as a progress-indexed interaction path $X(s)$ and a positive interval-time profile. The latter defines a monotone time law $t(s)$, yielding controller commands $X(s(t))$. For a given path, alternative executions are expressed through the time profile, allowing chunk-wise speed choices without changing the geometric prediction target. This representation supports a staged post-training procedure: demonstrations and DAgger interventions establish a target-domain prior, Speed-DQN learns execution multipliers from robot interaction, and Path-AWR uses rollout outcomes to refine the diffusion path generator. A path-conditioned action expert realizes the resulting motions while maintaining distinct learning interfaces for path generation and execution timing. Across three tasks, the complete method achieves $58/60$ successes versus $57/60$ for PathTime-VLA under BC + DAgger at fixed $1\\times$, with approximately $39$-$52\\%$ shorter mean completion times over successful trials."
    }
]