الملخص
تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.
We study model-free online reinforcement learning (RL) for episodic tabular Markov decision processes, focusing on both gap-dependent regret and policy switching cost. While fine-grained gap-dependent analysis has been established for model-free RL algorithms using Hoeffding-type exploration bonuses, such results for model-free algorithms with variance-based exploration bonuses remain unknown, despite their superior worst-case and coarse-grained gap-dependent guarantees. In this paper, we resolve this open problem by establishing the first fine-grained gap-dependent regret upper bound for UCB-Bernstein+, a refined UCB-Bernstein algorithm, in variance-aware model-free online RL. Moreover, by integrating a stage-wise policy update design into our fine-grained framework and using refined variance-based bonuses, we achieve the best-known gap-dependent local switching cost to date. In addition, our analysis yields improved worst-case guarantees for both regret and local switching cost over the original UCB-Bernstein algorithm. Numerical experiments further demonstrate that UCB-Bernstein+ achieves favorable empirical performance in both regret and local switching cost.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhang, H., Xue, L., & Zheng, Z. (2026). حدود مدركة للتباين ومعتمدة على الفجوة دقيقة الحبيبات للتعلم المعزز عبر الإنترنت. https://omanscience.com/ar/articles/variance-aware-fine-grained-gap-dependent-bounds-for-online-reinforcement-learning
MLA 9
Zhang, Haochen, et al. "حدود مدركة للتباين ومعتمدة على الفجوة دقيقة الحبيبات للتعلم المعزز عبر الإنترنت." https://omanscience.com/ar/articles/variance-aware-fine-grained-gap-dependent-bounds-for-online-reinforcement-learning.
شيكاغو (المؤلف–التاريخ)
Zhang, Haochen, Lingzhou Xue, and Zhong Zheng. 2026. "حدود مدركة للتباين ومعتمدة على الفجوة دقيقة الحبيبات للتعلم المعزز عبر الإنترنت." https://omanscience.com/ar/articles/variance-aware-fine-grained-gap-dependent-bounds-for-online-reinforcement-learning.
هارفارد
Zhang, H., Xue, L. and Zheng, Z. (2026) 'حدود مدركة للتباين ومعتمدة على الفجوة دقيقة الحبيبات للتعلم المعزز عبر الإنترنت', Available at: https://omanscience.com/ar/articles/variance-aware-fine-grained-gap-dependent-bounds-for-online-reinforcement-learning.
فانكوفر
Zhang H, Xue L, Zheng Z. حدود مدركة للتباين ومعتمدة على الفجوة دقيقة الحبيبات للتعلم المعزز عبر الإنترنت. https://omanscience.com/ar/articles/variance-aware-fine-grained-gap-dependent-bounds-for-online-reinforcement-learning
IEEE
H. Zhang, L. Xue, and Z. Zheng, "حدود مدركة للتباين ومعتمدة على الفجوة دقيقة الحبيبات للتعلم المعزز عبر الإنترنت," https://omanscience.com/ar/articles/variance-aware-fine-grained-gap-dependent-bounds-for-online-reinforcement-learning.