Abstract
Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and value vectors to write to the matrix-valued hidden state. We generalize this construction and propose triadic linear attention, which writes the triadic outer product of a key, a second key, and a value, into a third-order (i.e., 3D) tensor state, and reads from it by contracting both key axes with two queries. An $E$-dimensional second key thus yields an $E$-fold increase in state size while adding only two projections. Triadic linear attention is compatible with data-dependent forgetting, the delta rule, and chunkwise-parallel training. Applied to Gated DeltaNet and scalar-gated linear attention, triadic linear attention substantially improves long-context language modeling and recall, outperforming alternatives that enlarge the state.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Sieberling, O., Runwal, B., Jin, D., Chin, R., Panda, R., & Kim, Y. (2026). Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling. https://omanscience.com/en/articles/triadic-linear-attention-three-dimensional-recurrent-states-for-long-context-sequence-modeling
MLA 9
Sieberling, Oliver, et al. "Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling." https://omanscience.com/en/articles/triadic-linear-attention-three-dimensional-recurrent-states-for-long-context-sequence-modeling.
Chicago (author–date)
Sieberling, Oliver, Bharat Runwal, David Jin, Ryan Chin, Rameswar Panda, and Yoon Kim. 2026. "Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling." https://omanscience.com/en/articles/triadic-linear-attention-three-dimensional-recurrent-states-for-long-context-sequence-modeling.
Harvard
Sieberling, O., Runwal, B., Jin, D., Chin, R., Panda, R. and Kim, Y. (2026) 'Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling', Available at: https://omanscience.com/en/articles/triadic-linear-attention-three-dimensional-recurrent-states-for-long-context-sequence-modeling.
Vancouver
Sieberling O, Runwal B, Jin D, Chin R, Panda R, Kim Y. Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling. https://omanscience.com/en/articles/triadic-linear-attention-three-dimensional-recurrent-states-for-long-context-sequence-modeling
IEEE
O. Sieberling, B. Runwal, D. Jin, R. Chin, R. Panda, and Y. Kim, "Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling," https://omanscience.com/en/articles/triadic-linear-attention-three-dimensional-recurrent-states-for-long-context-sequence-modeling.