Abstract

Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Jun, Y., Choi, K., Kim, Y., Jin, S., Park, S., Park, J., & Ye, J. C. (2026). DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies. https://omanscience.com/en/articles/driftopd-sequence-level-reverse-kl-distillation-for-one-step-vla-policies

MLA 9

Jun, Youngjun, et al. "DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies." https://omanscience.com/en/articles/driftopd-sequence-level-reverse-kl-distillation-for-one-step-vla-policies.

Chicago (author–date)

Jun, Youngjun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, and Jong Chul Ye. 2026. "DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies." https://omanscience.com/en/articles/driftopd-sequence-level-reverse-kl-distillation-for-one-step-vla-policies.

Harvard

Jun, Y., Choi, K., Kim, Y., Jin, S., Park, S., Park, J. and Ye, J. C. (2026) 'DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies', Available at: https://omanscience.com/en/articles/driftopd-sequence-level-reverse-kl-distillation-for-one-step-vla-policies.

Vancouver

Jun Y, Choi K, Kim Y, Jin S, Park S, Park J, et al. DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies. https://omanscience.com/en/articles/driftopd-sequence-level-reverse-kl-distillation-for-one-step-vla-policies

IEEE

Y. Jun, K. Choi, Y. Kim, S. Jin, S. Park, J. Park, and J. C. Ye, "DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies," https://omanscience.com/en/articles/driftopd-sequence-level-reverse-kl-distillation-for-one-step-vla-policies.