Abstract
Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning. We introduce \textbf{FIND}, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. FIND reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem: instead of restoring a predefined scene after each rollout, it uses the resulting scene to determine what to practice next. A vision--language agent identifies feasible tasks from a predefined library, prioritizes those with lower recent success rates, and evaluates outcomes using paired pre- and post-execution observations. We instantiate FIND with a frozen $π_{0.5}$ VLA and residual off-policy RL. Across eight real-world manipulation tasks, the independent human-assessed success rate improves from $55\%$ to $71.9\%$. A representative run completes 456 autonomous episodes within 6 hours of interaction, requiring 30 scene-recovery interventions and no human-provided reward labels during online learning. Ablations and systematic evaluations further examine key design choices, agent evaluation accuracy, and human intervention requirements. Our website is made publicly available at: FIND.github.io.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Fang, Y., Li, Z., Tong, H., Liu, P., & Chalvatzaki, G. (2026). Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models. https://omanscience.com/en/articles/find-something-you-can-t-do-agentic-real-world-reinforcement-learning-for-self-improving-vla-models
MLA 9
Fang, Yuan, et al. "Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models." https://omanscience.com/en/articles/find-something-you-can-t-do-agentic-real-world-reinforcement-learning-for-self-improving-vla-models.
Chicago (author–date)
Fang, Yuan, Zechu Li, Haolei Tong, Puze Liu, and Georgia Chalvatzaki. 2026. "Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models." https://omanscience.com/en/articles/find-something-you-can-t-do-agentic-real-world-reinforcement-learning-for-self-improving-vla-models.
Harvard
Fang, Y., Li, Z., Tong, H., Liu, P. and Chalvatzaki, G. (2026) 'Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models', Available at: https://omanscience.com/en/articles/find-something-you-can-t-do-agentic-real-world-reinforcement-learning-for-self-improving-vla-models.
Vancouver
Fang Y, Li Z, Tong H, Liu P, Chalvatzaki G. Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models. https://omanscience.com/en/articles/find-something-you-can-t-do-agentic-real-world-reinforcement-learning-for-self-improving-vla-models
IEEE
Y. Fang, Z. Li, H. Tong, P. Liu, and G. Chalvatzaki, "Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models," https://omanscience.com/en/articles/find-something-you-can-t-do-agentic-real-world-reinforcement-learning-for-self-improving-vla-models.