Abstract

Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a single suboptimal mode, and optimization over the nonconvex Lagrangian landscape can be unstable. Diffusion and flow policies can represent such distributions, but recent work with a diffusion actor relies on estimating and matching the score of an augmented-Lagrangian target policy. Instead, we differentiate the augmented objective directly through the generation path of a flow policy, so no score needs to be estimated. Because a flow policy lacks a readily available action log-density for entropy regularization, we build on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and propose Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL. We formulate its update as a constrained one-ended generalized Schrödinger bridge and show that, for each source draw, this path-space problem is exactly an entropy-regularized problem in action space. At positive noise, its solution reweights the reward-only action distribution only where the estimated cost exceeds a threshold set by the Lagrange multiplier. As the noise vanishes, the optimal value converges to that of a least-energy map objective that the flow policy optimizes directly. Across seven Safety-Gymnasium tasks, RAFALE achieves competitive reward with mean final cost within budget on every task, whereas strong baselines trade one for the other; ablations support the necessity of both its augmented objective and its flow actor.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, B., Kim, M., & Herbert, S. (2026). Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View. https://omanscience.com/en/articles/constrained-flow-policy-updates-a-generalized-schr-dinger-bridge-view

MLA 9

Li, Boyang, et al. "Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View." https://omanscience.com/en/articles/constrained-flow-policy-updates-a-generalized-schr-dinger-bridge-view.

Chicago (author–date)

Li, Boyang, Matthew Kim, and Sylvia Herbert. 2026. "Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View." https://omanscience.com/en/articles/constrained-flow-policy-updates-a-generalized-schr-dinger-bridge-view.

Harvard

Li, B., Kim, M. and Herbert, S. (2026) 'Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View', Available at: https://omanscience.com/en/articles/constrained-flow-policy-updates-a-generalized-schr-dinger-bridge-view.

Vancouver

Li B, Kim M, Herbert S. Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View. https://omanscience.com/en/articles/constrained-flow-policy-updates-a-generalized-schr-dinger-bridge-view

IEEE

B. Li, M. Kim, and S. Herbert, "Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View," https://omanscience.com/en/articles/constrained-flow-policy-updates-a-generalized-schr-dinger-bridge-view.