Abstract
Simulation is a fundamental computational primitive in reinforcement learning (RL), yet conventional simulation explicitly generates individual trajectories even when downstream procedures use only aggregate statistics. To address this, we develop an exact fast-simulation framework for finite-horizon tabular Markov decision processes. Our framework has two complementary modes. In direct batch simulation, a batch is represented by its aggregate Markov flow. With sufficient parallel simulation resources, this flow can be obtained by trajectory aggregation; when such simulation is unavailable or costly but the initial state and transition distributions are directly accessible, we instead generate an identically distributed flow through forward Markov-flow sampling without materializing individual trajectories. The latter reduces the simulator-side computational dependence on batch size $m$ from $O(m)$ to $O(1)$. In adaptive batch simulation, when batch length is determined by a data-dependent condition, exact multivariate-hypergeometric splitting recursively refines a candidate Markov flow while preserving the conditional law, reducing the cost dependence on $m$ from $O(m)$ to $O(\log m)$. Together, these modes accelerate simulation by keeping trajectories aggregated whenever possible and refining flows only when required to locate data-dependent boundaries. The framework applies broadly across simulator-based, offline, and online batch or stage-based RL, as illustrated with representative algorithms from each setting.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhang, H., Xue, L., & Zheng, Z. (2026). Exact Fast Batch Simulation for Tabular Reinforcement Learning. https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning
MLA 9
Zhang, Haochen, et al. "Exact Fast Batch Simulation for Tabular Reinforcement Learning." https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.
Chicago (author–date)
Zhang, Haochen, Lingzhou Xue, and Zhong Zheng. 2026. "Exact Fast Batch Simulation for Tabular Reinforcement Learning." https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.
Harvard
Zhang, H., Xue, L. and Zheng, Z. (2026) 'Exact Fast Batch Simulation for Tabular Reinforcement Learning', Available at: https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.
Vancouver
Zhang H, Xue L, Zheng Z. Exact Fast Batch Simulation for Tabular Reinforcement Learning. https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning
IEEE
H. Zhang, L. Xue, and Z. Zheng, "Exact Fast Batch Simulation for Tabular Reinforcement Learning," https://omanscience.com/en/articles/exact-fast-batch-simulation-for-tabular-reinforcement-learning.