الملخص
Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's probability of ever departing from OTS is at most the chosen $α_E$, without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by $38.4\%$ on the registered suite but reduced it by $7.5\%$ on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Gulati, M., Wang, K., & Au, W. (2026). When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity. https://omanscience.com/ar/articles/when-may-a-bandit-leave-its-anchor-e-process-authorized-thompson-sampling-under-non-stationarity
MLA 9
Gulati, Mayand, et al. "When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity." https://omanscience.com/ar/articles/when-may-a-bandit-leave-its-anchor-e-process-authorized-thompson-sampling-under-non-stationarity.
شيكاغو (المؤلف–التاريخ)
Gulati, Mayand, Kerong Wang, and WeiChen Au. 2026. "When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity." https://omanscience.com/ar/articles/when-may-a-bandit-leave-its-anchor-e-process-authorized-thompson-sampling-under-non-stationarity.
هارفارد
Gulati, M., Wang, K. and Au, W. (2026) 'When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity', Available at: https://omanscience.com/ar/articles/when-may-a-bandit-leave-its-anchor-e-process-authorized-thompson-sampling-under-non-stationarity.
فانكوفر
Gulati M, Wang K, Au W. When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity. https://omanscience.com/ar/articles/when-may-a-bandit-leave-its-anchor-e-process-authorized-thompson-sampling-under-non-stationarity
IEEE
M. Gulati, K. Wang, and W. Au, "When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity," https://omanscience.com/ar/articles/when-may-a-bandit-leave-its-anchor-e-process-authorized-thompson-sampling-under-non-stationarity.