[
    {
        "id": "osp-19538",
        "type": "article-journal",
        "title": "The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs",
        "author": [
            {
                "family": "Aremu",
                "given": "Toluwani"
            },
            {
                "family": "Baser",
                "given": "Manit"
            },
            {
                "family": "Gurusamy",
                "given": "Mohan"
            },
            {
                "family": "Lukas",
                "given": "Nils"
            },
            {
                "family": "Divakaran",
                "given": "Dinil Mon"
            }
        ],
        "URL": "https://omanscience.com/en/articles/the-fragility-of-trigger-tag-mechanisms-for-misuse-detection-in-open-weight-llms",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \\emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as generating phishing contents. Although these mechanisms borrow from established techniques, their use for conditional misuse detection in open-weight LLMs is relatively new. Therefore, existing research works have not systematically studied the robustness of trigger-tag mechanisms under adversarial attacks. To close this gap, (i)~we formalize trigger-tags and distinguish \\emph{token-level trigger-tags}, which introduce watermark-inspired signals during decoding, from \\emph{weight-level trigger-tags}, which learn backdoor-inspired associations between target conditions and detectable model behavior. Furthermore, (ii)~we introduce \\Untag, a unified attack framework that organizes their mechanism-specific attack surfaces into a common taxonomy. We evaluate representative token-level and weight-level trigger-tags using phishing as a case study. We find that while trigger-tags may provide useful evidence in controlled settings, our attacks render the existing trigger-tag mechanisms to be entirely ineffective. Consequently, we argue that these mechanisms should not be treated as robust misuse detectors when attackers can transform outputs or modify open weights."
    }
]