الملخص

This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizing rewards, and adhering to the behavior regularization imposed by the offline dataset. To tackle this trilogy challenge, we propose Q-learning Penalized Transformer policy (QPT), a \emph{training--inference consistent} framework that bridges conditional sequence modeling with constraint-aware value estimation. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost, retaining strong behavior regularization. To inject explicit safety semantics during learning, we augment sequence-model training with a Q-shaped penalty using learned reward and cost Q-functions to favor high return under low constraint violation. At inference, the same Q-functions enforce the cost threshold and choose the highest-reward feasible action, closing the loop between training and deployment. We provide a principled analysis under stylized near-deterministic CMDPs, characterizing how Q-penalized conditional generation improve safety and performance. Empirically, QPT consistently outperforms strong safe offline RL baselines across 38 tasks on the DSRL benchmark, and exhibits robust zero-shot adaptation to different constraint thresholds.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Hu, S., Wang, P., Hu, J., Zhou, Q., Hu, A., Shen, L., Zhang, Y., & Tao, D. (2026). Q-learning Penalized Transformer for Safe Offline Reinforcement Learning. https://omanscience.com/ar/articles/q-learning-penalized-transformer-for-safe-offline-reinforcement-learning

MLA 9

Hu, Shengchao, et al. "Q-learning Penalized Transformer for Safe Offline Reinforcement Learning." https://omanscience.com/ar/articles/q-learning-penalized-transformer-for-safe-offline-reinforcement-learning.

شيكاغو (المؤلف–التاريخ)

Hu, Shengchao, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu, Li Shen, Ya Zhang, and Dacheng Tao. 2026. "Q-learning Penalized Transformer for Safe Offline Reinforcement Learning." https://omanscience.com/ar/articles/q-learning-penalized-transformer-for-safe-offline-reinforcement-learning.

هارفارد

Hu, S., Wang, P., Hu, J., Zhou, Q., Hu, A., Shen, L., Zhang, Y. and Tao, D. (2026) 'Q-learning Penalized Transformer for Safe Offline Reinforcement Learning', Available at: https://omanscience.com/ar/articles/q-learning-penalized-transformer-for-safe-offline-reinforcement-learning.

فانكوفر

Hu S, Wang P, Hu J, Zhou Q, Hu A, Shen L, et al. Q-learning Penalized Transformer for Safe Offline Reinforcement Learning. https://omanscience.com/ar/articles/q-learning-penalized-transformer-for-safe-offline-reinforcement-learning

IEEE

S. Hu, P. Wang, J. Hu, Q. Zhou, A. Hu, L. Shen, Y. Zhang, and D. Tao, "Q-learning Penalized Transformer for Safe Offline Reinforcement Learning," https://omanscience.com/ar/articles/q-learning-penalized-transformer-for-safe-offline-reinforcement-learning.