Preprint Open access
Readable Before Actionable: Causal Tracing of Indirect Prompt Injection
Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on …