Abstract
We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameters, multiple actions, and state- and action-dependent costs. For restless bandits (RBs), a well-studied special case of WCMDPs, prior work has developed policies that achieve an $O(1/\sqrt{N})$ optimality gap under general conditions, and has further identified conditions under which policies can achieve a better-than-$1/\sqrt{N}$ optimality gap. However, for general WCMDPs, no prior result achieves an optimality gap better than $1/\sqrt{N}$. In this paper, we identify conditions analogous to those for RBs under which a better-than-$1/\sqrt{N}$ optimality gap is achievable, and design a policy that attains an $O(1/N)$ optimality gap. Notably, unlike prior approaches based on generalizing priority orderings, our policy is not priority-based but rather is designed to induce locally linear mean-field dynamics.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Hong, Y., Zhang, X., Xie, Q., Chen, Y., & Wang, W. (2026). Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs. https://omanscience.com/en/articles/achieving-an-o-1-n-optimality-gap-in-average-reward-weakly-coupled-mdps
MLA 9
Hong, Yige, et al. "Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs." https://omanscience.com/en/articles/achieving-an-o-1-n-optimality-gap-in-average-reward-weakly-coupled-mdps.
Chicago (author–date)
Hong, Yige, Xiangcheng Zhang, Qiaomin Xie, Yudong Chen, and Weina Wang. 2026. "Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs." https://omanscience.com/en/articles/achieving-an-o-1-n-optimality-gap-in-average-reward-weakly-coupled-mdps.
Harvard
Hong, Y., Zhang, X., Xie, Q., Chen, Y. and Wang, W. (2026) 'Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs', Available at: https://omanscience.com/en/articles/achieving-an-o-1-n-optimality-gap-in-average-reward-weakly-coupled-mdps.
Vancouver
Hong Y, Zhang X, Xie Q, Chen Y, Wang W. Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs. https://omanscience.com/en/articles/achieving-an-o-1-n-optimality-gap-in-average-reward-weakly-coupled-mdps
IEEE
Y. Hong, X. Zhang, Q. Xie, Y. Chen, and W. Wang, "Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs," https://omanscience.com/en/articles/achieving-an-o-1-n-optimality-gap-in-average-reward-weakly-coupled-mdps.