الملخص
Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost. Naively reducing the integration budget can severely degrade control, since policies optimized under full-step execution are not explicitly constrained to remain reliable under coarse numerical integration. We refer to this mismatch as the few-step discretization gap. To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution. Reward-aware online Reflow rectifies student-induced transport paths during on-policy learning, making the resulting policy more robust to coarse integration. A frozen Gaussian PPO controller supplies complementary action-space supervision at full and intermediate integration budgets, while the deployed policy remains a single flow student executed with one Euler step. Across Unitree Go2, Boston Dynamics Spot, Unitree H1, and Unitree G1, RFPO consistently preserves full-step control performance under one-step execution, with one-step returns remaining within 2.4% of their corresponding 64-step values across both zero and random initialization. On Unitree Go2, one-step execution retains 98.5% of the 64-step reward while reducing onboard mean inference latency from 4.39 ms to 0.08 ms, yielding a 54.9x speedup. Real-robot experiments further validate stable one-step locomotion. Code: https://github.com/AIGeeksGroup/RFPO. Website: https://aigeeksgroup.github.io/RFPO.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Huang, T., Pan, L., Wang, H., Zhang, Z., Qian, S., Li, Y., Guo, Y., Shi, B., & Tang, H. (2026). RFPO: Rectified Flow Policy Optimization for Embodied Control. https://omanscience.com/ar/articles/rfpo-rectified-flow-policy-optimization-for-embodied-control
MLA 9
Huang, Ting, et al. "RFPO: Rectified Flow Policy Optimization for Embodied Control." https://omanscience.com/ar/articles/rfpo-rectified-flow-policy-optimization-for-embodied-control.
شيكاغو (المؤلف–التاريخ)
Huang, Ting, Lisiyu Pan, Haoyu Wang, Zeyu Zhang, Siyuan Qian, Yanjun Li, Yandong Guo, Boxin Shi, and Hao Tang. 2026. "RFPO: Rectified Flow Policy Optimization for Embodied Control." https://omanscience.com/ar/articles/rfpo-rectified-flow-policy-optimization-for-embodied-control.
هارفارد
Huang, T., Pan, L., Wang, H., Zhang, Z., Qian, S., Li, Y., Guo, Y., Shi, B. and Tang, H. (2026) 'RFPO: Rectified Flow Policy Optimization for Embodied Control', Available at: https://omanscience.com/ar/articles/rfpo-rectified-flow-policy-optimization-for-embodied-control.
فانكوفر
Huang T, Pan L, Wang H, Zhang Z, Qian S, Li Y, et al. RFPO: Rectified Flow Policy Optimization for Embodied Control. https://omanscience.com/ar/articles/rfpo-rectified-flow-policy-optimization-for-embodied-control
IEEE
T. Huang, L. Pan, H. Wang, Z. Zhang, S. Qian, Y. Li, Y. Guo, B. Shi, and H. Tang, "RFPO: Rectified Flow Policy Optimization for Embodied Control," https://omanscience.com/ar/articles/rfpo-rectified-flow-policy-optimization-for-embodied-control.