Abstract
Even a few action perturbations can substantially degrade the performance of a deployed decision policy. Certifying the resulting return loss is challenging in stochastic environments, where returns vary even without an attack. We introduce FoSeRL, a framework for certifying deployed RL policies against precommitted, temporally sparse action attacks. The deployed policy is unchanged, with no smoothing or retraining. Certification requires a resettable simulator supporting shared randomness and independent one-step successor queries, but no analytical dynamics model. FoSeRL certifies that an attacked episode loses no more than a prescribed amount of return relative to the same episode unattacked, with at least a target probability and at a user-specified confidence level. Both runs share the initial state and randomness, so the measured loss reflects the attack, not the episode; carrying the running return gap as a state coordinate makes it the terminal value, reducing trajectory-level certification to terminal safety. Time-dependent barrier conditions on the augmented state bound the terminal failure probability: satisfied exactly, they certify every admissible precommitted attack; learned from sampled trajectories and verified on held-out data, they certify the same guarantee under a specified attack-episode setting. Across six stochastic continuous-control environments and three RL policy families (TD3, SAC, and PPO), FoSeRL certifies non-trivial cardinality--magnitude robustness frontiers, achieves substantially larger certified budgets than policy smoothing, and reveals marked robustness differences among policies with comparable nominal performance.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Taheri, S., Ganguly, D. K., Křetínský, J., & Zamani, M. (2026). FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies. https://omanscience.com/en/articles/foserl-formal-sequential-robustness-certification-for-reinforcement-learning-policies
MLA 9
Taheri, Sara, et al. "FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies." https://omanscience.com/en/articles/foserl-formal-sequential-robustness-certification-for-reinforcement-learning-policies.
Chicago (author–date)
Taheri, Sara, Deep Kumar Ganguly, Jan Křetínský, and Majid Zamani. 2026. "FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies." https://omanscience.com/en/articles/foserl-formal-sequential-robustness-certification-for-reinforcement-learning-policies.
Harvard
Taheri, S., Ganguly, D. K., Křetínský, J. and Zamani, M. (2026) 'FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies', Available at: https://omanscience.com/en/articles/foserl-formal-sequential-robustness-certification-for-reinforcement-learning-policies.
Vancouver
Taheri S, Ganguly DK, Křetínský J, Zamani M. FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies. https://omanscience.com/en/articles/foserl-formal-sequential-robustness-certification-for-reinforcement-learning-policies
IEEE
S. Taheri, D. K. Ganguly, J. Křetínský, and M. Zamani, "FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies," https://omanscience.com/en/articles/foserl-formal-sequential-robustness-certification-for-reinforcement-learning-policies.