الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution once every $S \ge 1$ gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for policy gradient methods, and on-policy sampling (i.e., a small $S$, ideally $1$) is often viewed as crucial to their success; yet in prominent application like post-training large language models, reward-guided self-training has proved to be effective even when the rollout distribution is updated infrequently, but theoretical understanding remains limited for the convergence properties of these off-policy methods. To bridge these gaps, we develop a unified theory for RE(S) that covers the full spectrum of $S \ge 1$: it can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution. For multi-arm bandits with softmax policies, our in-depth analysis and numerical experiments reveal three key findings: (1) for any fixed $S$, RE(S) enjoys global convergence to the optimal policy as the number of rollout distribution updates $B = \lfloor T / S \rfloor \rightarrow \infty$, where $T$ denotes the number of gradient steps; (2) we prove tight two-sided bounds showing that the suboptimality gap of RE(S) achieves an asymptotic $Θ(1 / T)$ convergence rate, while $S$ only affects the length of a burn-in phase; (3) when initialized at a weak policy with a small optimal-action probability, RE(1) gets trapped around suboptimal policies for a long period, whereas RE(S) with a suitable $S$ avoids the detour and achieves significantly faster convergence to the global optimum, highlighting the benefits of off-policyness in this case.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wang, Z., Chen, Y., Li, Y., & Ding, B. (2026). الضبط الدقيق على البيانات المولَّدة ذاتيًا والموزونة بالمكافأة: ديناميكيات التعلم ومعدلات التقارب وفوائد خارج السياسة. https://omanscience.com/ar/articles/fine-tuning-on-self-generated-and-reward-weighted-data-learning-dynamics-convergence-rates-and-benefits-of-off-policyness

MLA 9

Wang, Zhiwei, et al. "الضبط الدقيق على البيانات المولَّدة ذاتيًا والموزونة بالمكافأة: ديناميكيات التعلم ومعدلات التقارب وفوائد خارج السياسة." https://omanscience.com/ar/articles/fine-tuning-on-self-generated-and-reward-weighted-data-learning-dynamics-convergence-rates-and-benefits-of-off-policyness.

شيكاغو (المؤلف–التاريخ)

Wang, Zhiwei, Yanxi Chen, Yaliang Li, and Bolin Ding. 2026. "الضبط الدقيق على البيانات المولَّدة ذاتيًا والموزونة بالمكافأة: ديناميكيات التعلم ومعدلات التقارب وفوائد خارج السياسة." https://omanscience.com/ar/articles/fine-tuning-on-self-generated-and-reward-weighted-data-learning-dynamics-convergence-rates-and-benefits-of-off-policyness.

هارفارد

Wang, Z., Chen, Y., Li, Y. and Ding, B. (2026) 'الضبط الدقيق على البيانات المولَّدة ذاتيًا والموزونة بالمكافأة: ديناميكيات التعلم ومعدلات التقارب وفوائد خارج السياسة', Available at: https://omanscience.com/ar/articles/fine-tuning-on-self-generated-and-reward-weighted-data-learning-dynamics-convergence-rates-and-benefits-of-off-policyness.

فانكوفر

Wang Z, Chen Y, Li Y, Ding B. الضبط الدقيق على البيانات المولَّدة ذاتيًا والموزونة بالمكافأة: ديناميكيات التعلم ومعدلات التقارب وفوائد خارج السياسة. https://omanscience.com/ar/articles/fine-tuning-on-self-generated-and-reward-weighted-data-learning-dynamics-convergence-rates-and-benefits-of-off-policyness

IEEE

Z. Wang, Y. Chen, Y. Li, and B. Ding, "الضبط الدقيق على البيانات المولَّدة ذاتيًا والموزونة بالمكافأة: ديناميكيات التعلم ومعدلات التقارب وفوائد خارج السياسة," https://omanscience.com/ar/articles/fine-tuning-on-self-generated-and-reward-weighted-data-learning-dynamics-convergence-rates-and-benefits-of-off-policyness.