Abstract

As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. We observe that even frontier reasoning models can fail to recover the current state. We find that in open-source models probes can still recover the updated value when the model answers with an old one, pointing to a failure to select information that remains available. Component tests in Qwen and Pythia identify a mechanism for this selection failure: attention drift, where attention favors old values over the current one when producing an answer. We study a one-layer transformer to mathematically understand how this phenomenon happens: when attention scores are similar, several old values can together receive more attention than the current value. Guided by this explanation, we redirect attention toward the current value without further training. When the current value is requested directly, adjusting this intervention for each input corrects most old-value errors across various model families while preserving nearly all initially correct answers. Reliable context management therefore requires more than remembering updated information: models must use it to guide their answers.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Guo, J., Fang, Y., Gu, S., Spanos, C., Demmel, J., & Lavaei, J. (2026). When Context Changes: Understanding Update Failures in LLMs. https://omanscience.com/en/articles/when-context-changes-understanding-update-failures-in-llms

MLA 9

Guo, Junyu, et al. "When Context Changes: Understanding Update Failures in LLMs." https://omanscience.com/en/articles/when-context-changes-understanding-update-failures-in-llms.

Chicago (author–date)

Guo, Junyu, Yuchen Fang, Shangding Gu, Costas Spanos, James Demmel, and Javad Lavaei. 2026. "When Context Changes: Understanding Update Failures in LLMs." https://omanscience.com/en/articles/when-context-changes-understanding-update-failures-in-llms.

Harvard

Guo, J., Fang, Y., Gu, S., Spanos, C., Demmel, J. and Lavaei, J. (2026) 'When Context Changes: Understanding Update Failures in LLMs', Available at: https://omanscience.com/en/articles/when-context-changes-understanding-update-failures-in-llms.

Vancouver

Guo J, Fang Y, Gu S, Spanos C, Demmel J, Lavaei J. When Context Changes: Understanding Update Failures in LLMs. https://omanscience.com/en/articles/when-context-changes-understanding-update-failures-in-llms

IEEE

J. Guo, Y. Fang, S. Gu, C. Spanos, J. Demmel, and J. Lavaei, "When Context Changes: Understanding Update Failures in LLMs," https://omanscience.com/en/articles/when-context-changes-understanding-update-failures-in-llms.