Abstract
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, $y=f(x^\top w)+\varepsilon $. Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of $f$, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Letey, M., Rysmakhanov, A., Lu, Y. M., Pehlevan, C., & Zavatone-Veth, J. (2026). What can linear attention learn from nonlinear teachers in-context? https://omanscience.com/en/articles/what-can-linear-attention-learn-from-nonlinear-teachers-in-context
MLA 9
Letey, Mary, et al. "What can linear attention learn from nonlinear teachers in-context?" https://omanscience.com/en/articles/what-can-linear-attention-learn-from-nonlinear-teachers-in-context.
Chicago (author–date)
Letey, Mary, Arman Rysmakhanov, Yue M. Lu, Cengiz Pehlevan, and Jacob Zavatone-Veth. 2026. "What can linear attention learn from nonlinear teachers in-context?" https://omanscience.com/en/articles/what-can-linear-attention-learn-from-nonlinear-teachers-in-context.
Harvard
Letey, M., Rysmakhanov, A., Lu, Y. M., Pehlevan, C. and Zavatone-Veth, J. (2026) 'What can linear attention learn from nonlinear teachers in-context?', Available at: https://omanscience.com/en/articles/what-can-linear-attention-learn-from-nonlinear-teachers-in-context.
Vancouver
Letey M, Rysmakhanov A, Lu YM, Pehlevan C, Zavatone-Veth J. What can linear attention learn from nonlinear teachers in-context? https://omanscience.com/en/articles/what-can-linear-attention-learn-from-nonlinear-teachers-in-context
IEEE
M. Letey, A. Rysmakhanov, Y. M. Lu, C. Pehlevan, and J. Zavatone-Veth, "What can linear attention learn from nonlinear teachers in-context?," https://omanscience.com/en/articles/what-can-linear-attention-learn-from-nonlinear-teachers-in-context.