Abstract

Simulation is a fundamental computational primitive in reinforcement learning (RL), yet conventional simulation explicitly generates individual trajectories even when downstream procedures use only aggregate statistics. To address this, we develop an exact fast-simulation framework for finite-horizon tabular Markov decision processes. Our framework has two complementary modes. In direct batch simulation, a batch is represented by its aggregate Markov flow. With sufficient parallel simulation resources, this flow can be obtained by trajectory aggregation; when such simulation is unavailable or costly but the initial state and transition distributions are directly accessible, we instead generate an identically distributed flow through forward Markov-flow sampling without materializing individual trajectories. The latter reduces the simulator-side computational dependence on batch size $m$ from $O(m)$ to $O(1)$. In adaptive batch simulation, when batch length is determined by a data-dependent condition, exact multivariate-hypergeometric splitting recursively refines a candidate Markov flow while preserving the conditional law, reducing the cost dependence on $m$ from $O(m)$ to $O(\log m)$. Together, these modes accelerate simulation by keeping trajectories aggregated whenever possible and refining flows only when required to locate data-dependent boundaries. The framework applies broadly across simulator-based, offline, and online batch or stage-based RL, as illustrated with representative algorithms from each setting.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, H., Xue, L., & Zheng, Z. (2026). Exact Fast Batch Simulation for Tabular Reinforcement Learning. https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning

MLA 9

Zhang, Haochen, et al. "Exact Fast Batch Simulation for Tabular Reinforcement Learning." https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.

Chicago (author–date)

Zhang, Haochen, Lingzhou Xue, and Zhong Zheng. 2026. "Exact Fast Batch Simulation for Tabular Reinforcement Learning." https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.

Harvard

Zhang, H., Xue, L. and Zheng, Z. (2026) 'Exact Fast Batch Simulation for Tabular Reinforcement Learning', Available at: https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.

Vancouver

Zhang H, Xue L, Zheng Z. Exact Fast Batch Simulation for Tabular Reinforcement Learning. https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning

IEEE

H. Zhang, L. Xue, and Z. Zheng, "Exact Fast Batch Simulation for Tabular Reinforcement Learning," https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.