الملخص
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhu, F. G., Xu, Q., Deng, Z., Hua, Z., Simon, L., Oh, J., & Simmons, R. (2026). Iterative Policy Refinement through Semantic Rollout Analysis. https://omanscience.com/ar/articles/iterative-policy-refinement-through-semantic-rollout-analysis
MLA 9
Zhu, Feiyu Gavin, et al. "Iterative Policy Refinement through Semantic Rollout Analysis." https://omanscience.com/ar/articles/iterative-policy-refinement-through-semantic-rollout-analysis.
شيكاغو (المؤلف–التاريخ)
Zhu, Feiyu Gavin, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, and Reid Simmons. 2026. "Iterative Policy Refinement through Semantic Rollout Analysis." https://omanscience.com/ar/articles/iterative-policy-refinement-through-semantic-rollout-analysis.
هارفارد
Zhu, F. G., Xu, Q., Deng, Z., Hua, Z., Simon, L., Oh, J. and Simmons, R. (2026) 'Iterative Policy Refinement through Semantic Rollout Analysis', Available at: https://omanscience.com/ar/articles/iterative-policy-refinement-through-semantic-rollout-analysis.
فانكوفر
Zhu FG, Xu Q, Deng Z, Hua Z, Simon L, Oh J, et al. Iterative Policy Refinement through Semantic Rollout Analysis. https://omanscience.com/ar/articles/iterative-policy-refinement-through-semantic-rollout-analysis
IEEE
F. G. Zhu, Q. Xu, Z. Deng, Z. Hua, L. Simon, J. Oh, and R. Simmons, "Iterative Policy Refinement through Semantic Rollout Analysis," https://omanscience.com/ar/articles/iterative-policy-refinement-through-semantic-rollout-analysis.