Abstract

Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated. We ask whether an agent can infer the monitor's decision rule from this feedback and then selectively edit its activations to evade detection. Unlike prior evasion attacks, the model is never explicitly told what the monitor detects. Surprisingly, off-the-shelf models already produce activation edits aligned with the monitored direction, but at insufficient magnitude for evasion. Simply scaling up these edits by a factor of 8 reduces the monitor's TPR from 100% to 27%. A rank-1 LoRA amplifies this behavior into effective evasion within the forward pass, reducing TPR further to 4% on held-out concept monitors while leaving other concepts at their normal detection rates. Capabilities on standard benchmarks are retained under this finetuning, and the evasion skill survives retraining the monitors on the new activations. Mechanistically, we find evidence that the model computes its activation edit from the prior in-context turns, and show that the edit becomes more aligned with the monitored direction as more examples are provided. These results demonstrate feedback-conditioned control over activations and suggest that latent monitoring should be treated as an interactive process in which agents can observe and respond to oversight measures.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Keenan, H. L., Leckie, C., & Erfani, S. (2026). LLMs Learn to Evade Latent Monitors from Prior Feedback Alone. https://omanscience.com/en/articles/llms-learn-to-evade-latent-monitors-from-prior-feedback-alone

MLA 9

Keenan, Hugo Lyons, et al. "LLMs Learn to Evade Latent Monitors from Prior Feedback Alone." https://omanscience.com/en/articles/llms-learn-to-evade-latent-monitors-from-prior-feedback-alone.

Chicago (author–date)

Keenan, Hugo Lyons, Christopher Leckie, and Sarah Erfani. 2026. "LLMs Learn to Evade Latent Monitors from Prior Feedback Alone." https://omanscience.com/en/articles/llms-learn-to-evade-latent-monitors-from-prior-feedback-alone.

Harvard

Keenan, H. L., Leckie, C. and Erfani, S. (2026) 'LLMs Learn to Evade Latent Monitors from Prior Feedback Alone', Available at: https://omanscience.com/en/articles/llms-learn-to-evade-latent-monitors-from-prior-feedback-alone.

Vancouver

Keenan HL, Leckie C, Erfani S. LLMs Learn to Evade Latent Monitors from Prior Feedback Alone. https://omanscience.com/en/articles/llms-learn-to-evade-latent-monitors-from-prior-feedback-alone

IEEE

H. L. Keenan, C. Leckie, and S. Erfani, "LLMs Learn to Evade Latent Monitors from Prior Feedback Alone," https://omanscience.com/en/articles/llms-learn-to-evade-latent-monitors-from-prior-feedback-alone.