Abstract

Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model. At test time, each output dimension dynamically selects among multiple linear attention components based on the input. Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lee, H. D., Gonzalez, X., Zucchet, N., Buchanan, E. K., Fox, E. B., & Linderman, S. W. (2026). Switching Linear Attention. https://omanscience.com/en/articles/switching-linear-attention

MLA 9

Lee, Hyun Dong, et al. "Switching Linear Attention." https://omanscience.com/en/articles/switching-linear-attention.

Chicago (author–date)

Lee, Hyun Dong, Xavier Gonzalez, Nicolas Zucchet, E. Kelly Buchanan, Emily B. Fox, and Scott W. Linderman. 2026. "Switching Linear Attention." https://omanscience.com/en/articles/switching-linear-attention.

Harvard

Lee, H. D., Gonzalez, X., Zucchet, N., Buchanan, E. K., Fox, E. B. and Linderman, S. W. (2026) 'Switching Linear Attention', Available at: https://omanscience.com/en/articles/switching-linear-attention.

Vancouver

Lee HD, Gonzalez X, Zucchet N, Buchanan EK, Fox EB, Linderman SW. Switching Linear Attention. https://omanscience.com/en/articles/switching-linear-attention

IEEE

H. D. Lee, X. Gonzalez, N. Zucchet, E. K. Buchanan, E. B. Fox, and S. W. Linderman, "Switching Linear Attention," https://omanscience.com/en/articles/switching-linear-attention.