Abstract
Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize and arbitrary finite critic initialization. The proof combines a resolvent-based auxiliary distribution with a decaying Bellman-violation correction and a potential weighted by inverse coordinate weights. We then study stochastic TD-PMD with general strongly convex mirror maps under a single off-policy Markov trajectory. With suitably chosen constant stepsizes and a finite-batch TD update, the method achieves an expected value gap of $ε$ after $\widetilde{O}(1/((1-γ)^5 \widetildeσ_b ε))$ transitions. The stochastic analysis relies on the trajectory-wise Lipschitz continuity of the regularizer, derived from uniform bounds on vertex Bregman divergences, together with a visitation-weighted resolvent estimate for signed critic-error propagation that yields an inverse-linear dependence on behavior coverage $\widetildeσ_b$. In contrast to many prior guarantees for regularized policy optimization, our sample-complexity guarantee holds without trajectory resets, generative-model access, or nested policy-evaluation loops. Numerical results are consistent with the theoretical convergence analysis.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chen, Q., Li, W., Sun, Y., & Wei, K. (2026). Fast Regularized Policy Mirror Descent with One-Step TD Updates. https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates
MLA 9
Chen, Qipei, et al. "Fast Regularized Policy Mirror Descent with One-Step TD Updates." https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.
Chicago (author–date)
Chen, Qipei, Wenye Li, Yule Sun, and Ke Wei. 2026. "Fast Regularized Policy Mirror Descent with One-Step TD Updates." https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.
Harvard
Chen, Q., Li, W., Sun, Y. and Wei, K. (2026) 'Fast Regularized Policy Mirror Descent with One-Step TD Updates', Available at: https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.
Vancouver
Chen Q, Li W, Sun Y, Wei K. Fast Regularized Policy Mirror Descent with One-Step TD Updates. https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates
IEEE
Q. Chen, W. Li, Y. Sun, and K. Wei, "Fast Regularized Policy Mirror Descent with One-Step TD Updates," https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.