Abstract

A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yu, Z., Xing, W., Wei, Y., Yang, B., Ye, C., Li, G., & Han, M. (2026). The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models. https://omanscience.com/en/articles/the-attribution-blind-spot-layerwise-trajectory-diagnostics-for-source-reliance-in-retrieval-augmented-language-models

MLA 9

Yu, Zhe, et al. "The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models." https://omanscience.com/en/articles/the-attribution-blind-spot-layerwise-trajectory-diagnostics-for-source-reliance-in-retrieval-augmented-language-models.

Chicago (author–date)

Yu, Zhe, Wenpeng Xing, Yunzhao Wei, Bo Yang, Chen Ye, Gaolei Li, and Meng Han. 2026. "The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models." https://omanscience.com/en/articles/the-attribution-blind-spot-layerwise-trajectory-diagnostics-for-source-reliance-in-retrieval-augmented-language-models.

Harvard

Yu, Z., Xing, W., Wei, Y., Yang, B., Ye, C., Li, G. and Han, M. (2026) 'The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models', Available at: https://omanscience.com/en/articles/the-attribution-blind-spot-layerwise-trajectory-diagnostics-for-source-reliance-in-retrieval-augmented-language-models.

Vancouver

Yu Z, Xing W, Wei Y, Yang B, Ye C, Li G, et al. The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models. https://omanscience.com/en/articles/the-attribution-blind-spot-layerwise-trajectory-diagnostics-for-source-reliance-in-retrieval-augmented-language-models

IEEE

Z. Yu, W. Xing, Y. Wei, B. Yang, C. Ye, G. Li, and M. Han, "The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models," https://omanscience.com/en/articles/the-attribution-blind-spot-layerwise-trajectory-diagnostics-for-source-reliance-in-retrieval-augmented-language-models.