الملخص
Visuomotor policies can execute familiar tasks yet lack the corrective behavior needed after their own mistakes. We present a framework for recursive self-improvement through local recovery supervision. Each round audits the current policy, generates corrective demonstrations at supported failure states, and uses them to update the policy that drives the next round of collection. An offline auditor locates unresolved failures using coarse and dense temporal evidence and specifies observable repair goals. A fixed multimodal agent acts as a tool-using teacher, generating recovery actions through observation, computation, execution, and feedback. The frozen student tests whether each teacher endpoint supports further progress. If continuation fails, the system restores that endpoint and extends the demonstration. Action-level quality assessment then defines continuous training windows with aligned observations, quality weights, and validity masks. Only the student is deployed. In a preliminary LIBERO-Goal study, recovery-augmented post-training achieves 88 successful episodes out of 100 validation scenes, compared with 78 for original-data continuation from the same $π_0$ checkpoint. An earlier BC-RNN study on robomimic Can improves success from 102/130 to 112/130 using 26 local recovery segments. Both comparisons match 2,000 additional optimization steps.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhang, Y., Liu, X., & Zhang, Y. (2026). Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision. https://omanscience.com/ar/articles/recursive-self-improvement-of-visuomotor-policies-through-local-recovery-supervision
MLA 9
Zhang, Yuzhi, et al. "Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision." https://omanscience.com/ar/articles/recursive-self-improvement-of-visuomotor-policies-through-local-recovery-supervision.
شيكاغو (المؤلف–التاريخ)
Zhang, Yuzhi, Xinyu Liu, and Yu Zhang. 2026. "Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision." https://omanscience.com/ar/articles/recursive-self-improvement-of-visuomotor-policies-through-local-recovery-supervision.
هارفارد
Zhang, Y., Liu, X. and Zhang, Y. (2026) 'Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision', Available at: https://omanscience.com/ar/articles/recursive-self-improvement-of-visuomotor-policies-through-local-recovery-supervision.
فانكوفر
Zhang Y, Liu X, Zhang Y. Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision. https://omanscience.com/ar/articles/recursive-self-improvement-of-visuomotor-policies-through-local-recovery-supervision
IEEE
Y. Zhang, X. Liu, and Y. Zhang, "Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision," https://omanscience.com/ar/articles/recursive-self-improvement-of-visuomotor-policies-through-local-recovery-supervision.