Abstract
Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace. This coupling complicates adaptation from teleoperation: useful geometric guidance comes with timing shaped by interface delays and operator behavior. Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA. We introduce PathTime-VLA, which represents motion as a progress-indexed interaction path $X(s)$ and a positive interval-time profile. The latter defines a monotone time law $t(s)$, yielding controller commands $X(s(t))$. For a given path, alternative executions are expressed through the time profile, allowing chunk-wise speed choices without changing the geometric prediction target. This representation supports a staged post-training procedure: demonstrations and DAgger interventions establish a target-domain prior, Speed-DQN learns execution multipliers from robot interaction, and Path-AWR uses rollout outcomes to refine the diffusion path generator. A path-conditioned action expert realizes the resulting motions while maintaining distinct learning interfaces for path generation and execution timing. Across three tasks, the complete method achieves $58/60$ successes versus $57/60$ for PathTime-VLA under BC + DAgger at fixed $1\times$, with approximately $39$-$52\%$ shorter mean completion times over successful trials.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Huang, Q., Yang, Y., Zou, Z., Chen, A., Zhu, Z., Wei, Y., Xiong, R., & Wang, Y. (2026). PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies. https://omanscience.com/en/articles/pathtime-vla-path-time-decoupling-for-factorized-post-training-of-vision-language-action-policies
MLA 9
Huang, Qing, et al. "PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies." https://omanscience.com/en/articles/pathtime-vla-path-time-decoupling-for-factorized-post-training-of-vision-language-action-policies.
Chicago (author–date)
Huang, Qing, Yifei Yang, Ziqing Zou, Anzhe Chen, Zhenjie Zhu, Yufei Wei, Rong Xiong, and Yue Wang. 2026. "PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies." https://omanscience.com/en/articles/pathtime-vla-path-time-decoupling-for-factorized-post-training-of-vision-language-action-policies.
Harvard
Huang, Q., Yang, Y., Zou, Z., Chen, A., Zhu, Z., Wei, Y., Xiong, R. and Wang, Y. (2026) 'PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies', Available at: https://omanscience.com/en/articles/pathtime-vla-path-time-decoupling-for-factorized-post-training-of-vision-language-action-policies.
Vancouver
Huang Q, Yang Y, Zou Z, Chen A, Zhu Z, Wei Y, et al. PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies. https://omanscience.com/en/articles/pathtime-vla-path-time-decoupling-for-factorized-post-training-of-vision-language-action-policies
IEEE
Q. Huang, Y. Yang, Z. Zou, A. Chen, Z. Zhu, Y. Wei, R. Xiong, and Y. Wang, "PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies," https://omanscience.com/en/articles/pathtime-vla-path-time-decoupling-for-factorized-post-training-of-vision-language-action-policies.