Abstract
Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Dhouib, M., Elliker, C., Canesse, A., Jenny, M., Thil, L. A., El Sharkawy, M., Vanier, S., & Bursztein, E. (2026). RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents. https://omanscience.com/en/articles/raised-self-distillation-for-robustness-to-prompt-injection-in-llm-agents
MLA 9
Dhouib, Mohamed, et al. "RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents." https://omanscience.com/en/articles/raised-self-distillation-for-robustness-to-prompt-injection-in-llm-agents.
Chicago (author–date)
Dhouib, Mohamed, Clement Elliker, Alexi Canesse, Maël Jenny, Lucas-Andrei Thil, Mahammed El Sharkawy, Sonia Vanier, and Elie Bursztein. 2026. "RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents." https://omanscience.com/en/articles/raised-self-distillation-for-robustness-to-prompt-injection-in-llm-agents.
Harvard
Dhouib, M., Elliker, C., Canesse, A., Jenny, M., Thil, L. A., El Sharkawy, M., Vanier, S. and Bursztein, E. (2026) 'RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents', Available at: https://omanscience.com/en/articles/raised-self-distillation-for-robustness-to-prompt-injection-in-llm-agents.
Vancouver
Dhouib M, Elliker C, Canesse A, Jenny M, Thil LA, El Sharkawy M, et al. RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents. https://omanscience.com/en/articles/raised-self-distillation-for-robustness-to-prompt-injection-in-llm-agents
IEEE
M. Dhouib, C. Elliker, A. Canesse, M. Jenny, L. A. Thil, M. El Sharkawy, S. Vanier, and E. Bursztein, "RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents," https://omanscience.com/en/articles/raised-self-distillation-for-robustness-to-prompt-injection-in-llm-agents.