Abstract
Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progress and use it as dense feedback for policy improvement. Despite their architectural differences, existing progress-aware methods commonly formulate task progress as an explicit scalar prediction, providing limited structure for modeling how intermediate observations relate to the task goal, which may hinder effective transition-level credit assignment. We introduce Progress Field Reinforcement Learning (PF-RL), which learns a structured goal-conditioned progress representation over pretrained VLA features and converts it into dense credit for policy optimization. A lightweight shared Progress Field head maps current and goal representations into a compact progress space, where geometric distance induces goal-conditioned value, while complementary temporal and goal-structure objectives shape the learned geometry. Transition-level value changes naturally yield dense progress advantages, enabling fine-grained credit assignment for both offline policy improvement and online reinforcement fine-tuning. Extensive experiments on LIBERO, RoboTwin2.0, and real-world bimanual manipulation tasks show that PF-RL consistently improves policy performance over strong supervised fine-tuning, reinforcement fine-tuning, and progress-aware baselines.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Qing, Y., Kong, Y., Lin, S., Zhou, M., Fei, Y., Luo, S., Chi, Y., Gu, H., Liu, J., Wei, C., Hou, Z., & Zou, C. (2026). PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models. https://omanscience.com/en/articles/pf-rl-progress-field-reinforcement-learning-via-goal-conditioned-value-geometry-for-vision-language-action-models
MLA 9
Qing, Yunpeng, et al. "PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models." https://omanscience.com/en/articles/pf-rl-progress-field-reinforcement-learning-via-goal-conditioned-value-geometry-for-vision-language-action-models.
Chicago (author–date)
Qing, Yunpeng, Yilun Kong, Sixu Lin, Ming Zhou, Yiming Fei, Shuang Luo, Yixiao Chi, Haoming Gu, Jingyuan Liu, Changxu Wei, Zhi Hou, and Changqing Zou. 2026. "PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models." https://omanscience.com/en/articles/pf-rl-progress-field-reinforcement-learning-via-goal-conditioned-value-geometry-for-vision-language-action-models.
Harvard
Qing, Y., Kong, Y., Lin, S., Zhou, M., Fei, Y., Luo, S., Chi, Y., Gu, H., Liu, J., Wei, C., Hou, Z. and Zou, C. (2026) 'PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models', Available at: https://omanscience.com/en/articles/pf-rl-progress-field-reinforcement-learning-via-goal-conditioned-value-geometry-for-vision-language-action-models.
Vancouver
Qing Y, Kong Y, Lin S, Zhou M, Fei Y, Luo S, et al. PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models. https://omanscience.com/en/articles/pf-rl-progress-field-reinforcement-learning-via-goal-conditioned-value-geometry-for-vision-language-action-models
IEEE
Y. Qing, Y. Kong, S. Lin, M. Zhou, Y. Fei, S. Luo, Y. Chi, H. Gu, J. Liu, C. Wei, Z. Hou, and C. Zou, "PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models," https://omanscience.com/en/articles/pf-rl-progress-field-reinforcement-learning-via-goal-conditioned-value-geometry-for-vision-language-action-models.