Abstract
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation--attacker-chosen arguments in otherwise legitimate calls--is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Russinovich, M. (2026). CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering. https://omanscience.com/en/articles/countersteer-suppressing-indirect-prompt-injection-with-activation-steering
MLA 9
Russinovich, Mark. "CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering." https://omanscience.com/en/articles/countersteer-suppressing-indirect-prompt-injection-with-activation-steering.
Chicago (author–date)
Russinovich, Mark. 2026. "CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering." https://omanscience.com/en/articles/countersteer-suppressing-indirect-prompt-injection-with-activation-steering.
Harvard
Russinovich, M. (2026) 'CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering', Available at: https://omanscience.com/en/articles/countersteer-suppressing-indirect-prompt-injection-with-activation-steering.
Vancouver
Russinovich M. CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering. https://omanscience.com/en/articles/countersteer-suppressing-indirect-prompt-injection-with-activation-steering
IEEE
M. Russinovich, "CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering," https://omanscience.com/en/articles/countersteer-suppressing-indirect-prompt-injection-with-activation-steering.