[
    {
        "id": "osp-19903",
        "type": "article-journal",
        "title": "Empty Commitments: When Agents Promise What They Cannot Deliver",
        "author": [
            {
                "family": "Tang",
                "given": "Jiaqi"
            },
            {
                "family": "Shen",
                "given": "Bingyu"
            },
            {
                "family": "Wei",
                "given": "Lan"
            },
            {
                "family": "Lu",
                "given": "Qing"
            },
            {
                "family": "Ololade",
                "given": "Bethel"
            },
            {
                "family": "Galvis",
                "given": "Danny"
            },
            {
                "family": "Li",
                "given": "Miles Q."
            },
            {
                "family": "Hu",
                "given": "Bin"
            },
            {
                "family": "Li",
                "given": "Boyang"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/empty-commitments-when-agents-promise-what-they-cannot-deliver",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "A chatbot that says \"I will remind you tomorrow\" will not run again until the user writes. We call such a promise an empty commitment: a promise of action after the current turn that nothing in the agent's tools or runtime can carry out. Unlike a broken promise, its emptiness is decided by the agent's configuration at the moment of speaking, so it can be detected from a single turn, before deployment or at run time. We define empty commitments on top of commitment semantics, with three failure types, an anchoring condition for promises that a tool could make real, and an outcome taxonomy that separates these failures from honest deferrals and from over-refusal. We build a checker (a setup-blind detector, deterministic feasibility rules, and a response judge) validated against 400 human labels. On a controlled benchmark of 293 follow-up requests across five setups that add one persistence affordance at a time, four open-weight models of 8-14B parameters fail on 45.9% of responses when no tool exists and nothing is stated. A frontier model fails on 4.4%, but it gets there by deferring and asking, not by using the tools it has: promises made without the enabling call remain in every model. Telling the model its runtime, the cheapest fix, cuts open-weight failures nearly in half where nothing is doable and changes nothing where a scheduler exists; a directive capability card removes most failures at the largest cost in over-refusal; running the checker in the loop and rewriting flagged replies removes more at a smaller cost. Code, prompts, model outputs, and human labels are released."
    }
]