Abstract

Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize and arbitrary finite critic initialization. The proof combines a resolvent-based auxiliary distribution with a decaying Bellman-violation correction and a potential weighted by inverse coordinate weights. We then study stochastic TD-PMD with general strongly convex mirror maps under a single off-policy Markov trajectory. With suitably chosen constant stepsizes and a finite-batch TD update, the method achieves an expected value gap of $ε$ after $\widetilde{O}(1/((1-γ)^5 \widetildeσ_b ε))$ transitions. The stochastic analysis relies on the trajectory-wise Lipschitz continuity of the regularizer, derived from uniform bounds on vertex Bregman divergences, together with a visitation-weighted resolvent estimate for signed critic-error propagation that yields an inverse-linear dependence on behavior coverage $\widetildeσ_b$. In contrast to many prior guarantees for regularized policy optimization, our sample-complexity guarantee holds without trajectory resets, generative-model access, or nested policy-evaluation loops. Numerical results are consistent with the theoretical convergence analysis.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chen, Q., Li, W., Sun, Y., & Wei, K. (2026). Fast Regularized Policy Mirror Descent with One-Step TD Updates. https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates

MLA 9

Chen, Qipei, et al. "Fast Regularized Policy Mirror Descent with One-Step TD Updates." https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.

Chicago (author–date)

Chen, Qipei, Wenye Li, Yule Sun, and Ke Wei. 2026. "Fast Regularized Policy Mirror Descent with One-Step TD Updates." https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.

Harvard

Chen, Q., Li, W., Sun, Y. and Wei, K. (2026) 'Fast Regularized Policy Mirror Descent with One-Step TD Updates', Available at: https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.

Vancouver

Chen Q, Li W, Sun Y, Wei K. Fast Regularized Policy Mirror Descent with One-Step TD Updates. https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates

IEEE

Q. Chen, W. Li, Y. Sun, and K. Wei, "Fast Regularized Policy Mirror Descent with One-Step TD Updates," https://omanscience.com/en/articles/fast-regularized-policy-mirror-descent-with-one-step-td-updates.