Abstract

Actor-critic methods reuse past experience to improve sample efficiency. However, historical data are typically regarded as off-policy samples for the current policy-improvement update. This work introduces Regularized Dual Averaging Actor Critic (RDA2C), which assigns a distinct role to replay. In regularized dual averaging, the subsequent policy is determined by accumulated policy-improvement feedback, so historical advantage estimates contribute directly to the actor objective rather than solely to the most recent update. RDA2C stores state-action samples with critic-estimated advantage labels, fits a dual score model $Z_θ$ to the aggregated dataset, and derives the current policy from the accumulated score model using the entropy mirror map. In this way, replay defines an empirical dual objective from which the policy is computed. To analyze RDA2C, we establish a finite-time value-gap decomposition, separating the regularized dual-averaging term from errors due to stale-replay supervised fitting, critic bias, finite-buffer variance, and replay coverage, and stating the assumptions under which each error is bounded. RDA2C accepts advantage labels from any critic. With GAE labels, RDA2C outperforms PPO on six of eight MuJoCo tasks and eight of twelve Atari games. RDA2C also outperforms AAPDA, the closest dual-averaging baseline, on six of eight MuJoCo tasks. With twin-$Q$ labels, RDA2C matches SAC at matched batch size and update frequency.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Peng, N., Gordon, G. J., & Brantley, K. (2026). Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression. https://omanscience.com/en/articles/policy-as-data-replay-based-policy-dual-averaging-via-advantage-regression

MLA 9

Peng, Nianli, et al. "Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression." https://omanscience.com/en/articles/policy-as-data-replay-based-policy-dual-averaging-via-advantage-regression.

Chicago (author–date)

Peng, Nianli, Geoffrey J. Gordon, and Kianté Brantley. 2026. "Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression." https://omanscience.com/en/articles/policy-as-data-replay-based-policy-dual-averaging-via-advantage-regression.

Harvard

Peng, N., Gordon, G. J. and Brantley, K. (2026) 'Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression', Available at: https://omanscience.com/en/articles/policy-as-data-replay-based-policy-dual-averaging-via-advantage-regression.

Vancouver

Peng N, Gordon GJ, Brantley K. Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression. https://omanscience.com/en/articles/policy-as-data-replay-based-policy-dual-averaging-via-advantage-regression

IEEE

N. Peng, G. J. Gordon, and K. Brantley, "Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression," https://omanscience.com/en/articles/policy-as-data-replay-based-policy-dual-averaging-via-advantage-regression.