[
    {
        "id": "osp-22278",
        "type": "article-journal",
        "title": "Q-learning Penalized Transformer for Safe Offline Reinforcement Learning",
        "author": [
            {
                "family": "Hu",
                "given": "Shengchao"
            },
            {
                "family": "Wang",
                "given": "Peng"
            },
            {
                "family": "Hu",
                "given": "Jifeng"
            },
            {
                "family": "Zhou",
                "given": "Qiyang"
            },
            {
                "family": "Hu",
                "given": "Anning"
            },
            {
                "family": "Shen",
                "given": "Li"
            },
            {
                "family": "Zhang",
                "given": "Ya"
            },
            {
                "family": "Tao",
                "given": "Dacheng"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/q-learning-penalized-transformer-for-safe-offline-reinforcement-learning",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizing rewards, and adhering to the behavior regularization imposed by the offline dataset. To tackle this trilogy challenge, we propose Q-learning Penalized Transformer policy (QPT), a \\emph{training--inference consistent} framework that bridges conditional sequence modeling with constraint-aware value estimation. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost, retaining strong behavior regularization. To inject explicit safety semantics during learning, we augment sequence-model training with a Q-shaped penalty using learned reward and cost Q-functions to favor high return under low constraint violation. At inference, the same Q-functions enforce the cost threshold and choose the highest-reward feasible action, closing the loop between training and deployment. We provide a principled analysis under stylized near-deterministic CMDPs, characterizing how Q-penalized conditional generation improve safety and performance. Empirically, QPT consistently outperforms strong safe offline RL baselines across 38 tasks on the DSRL benchmark, and exhibits robust zero-shot adaptation to different constraint thresholds."
    }
]