Abstract
We develop an amortized, grid-free implementation of continuous Langevin dynamics based policy-value iteration for entropy-regularized, infinite-horizon relaxed stochastic control problems. The improvement rate of the exact iteration is a discounted aggregate of relative Fisher information between the policy and the Gibbs law of its Hamiltonian. The associated score residual is the velocity with which the control's Langevin dynamics transport its law. We project this velocity onto a conditional sampler shared across states, instead of one Langevin dynamics per state, and the value dynamics onto a parametric critic, estimating both projections at sampled states to obtain coupled actor--critic flows. The score loss measures the actor's agreement with the current critic, while the policy-evaluation residual measures the critic's agreement with the actor. We also derive gradient and Hessian residuals, including a Feynman--Kac representation for the gradient equation, to control errors not detected by the projected value iteration. An exact decomposition of the HJB residual combines these errors into a policy-suboptimality bound under verification and logarithmic Sobolev assumptions. In the linear-quadratic class, both projections are exact and recover the pointwise iteration, and we provide numerical experiments on general models to demonstrate the coupled actor--critic learning in high-dimensions.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Feng, Q., & Wang, G. (2026). Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems. https://omanscience.com/en/articles/amortized-score-hamiltonian-policy-iteration-a-grid-free-scheme-for-relaxed-stochastic-control-problems
MLA 9
Feng, Qi, and Gu Wang. "Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems." https://omanscience.com/en/articles/amortized-score-hamiltonian-policy-iteration-a-grid-free-scheme-for-relaxed-stochastic-control-problems.
Chicago (author–date)
Feng, Qi, and Gu Wang. 2026. "Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems." https://omanscience.com/en/articles/amortized-score-hamiltonian-policy-iteration-a-grid-free-scheme-for-relaxed-stochastic-control-problems.
Harvard
Feng, Q. and Wang, G. (2026) 'Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems', Available at: https://omanscience.com/en/articles/amortized-score-hamiltonian-policy-iteration-a-grid-free-scheme-for-relaxed-stochastic-control-problems.
Vancouver
Feng Q, Wang G. Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems. https://omanscience.com/en/articles/amortized-score-hamiltonian-policy-iteration-a-grid-free-scheme-for-relaxed-stochastic-control-problems
IEEE
Q. Feng, and G. Wang, "Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems," https://omanscience.com/en/articles/amortized-score-hamiltonian-policy-iteration-a-grid-free-scheme-for-relaxed-stochastic-control-problems.