Abstract
Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, Y., Han, D., Fu, J., & Huang, G. (2026). DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory. https://omanscience.com/en/articles/deltattt-layerwise-optimization-for-nonlinear-recurrent-memory
MLA 9
Li, Yining, et al. "DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory." https://omanscience.com/en/articles/deltattt-layerwise-optimization-for-nonlinear-recurrent-memory.
Chicago (author–date)
Li, Yining, Dongchen Han, Jie Fu, and Gao Huang. 2026. "DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory." https://omanscience.com/en/articles/deltattt-layerwise-optimization-for-nonlinear-recurrent-memory.
Harvard
Li, Y., Han, D., Fu, J. and Huang, G. (2026) 'DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory', Available at: https://omanscience.com/en/articles/deltattt-layerwise-optimization-for-nonlinear-recurrent-memory.
Vancouver
Li Y, Han D, Fu J, Huang G. DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory. https://omanscience.com/en/articles/deltattt-layerwise-optimization-for-nonlinear-recurrent-memory
IEEE
Y. Li, D. Han, J. Fu, and G. Huang, "DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory," https://omanscience.com/en/articles/deltattt-layerwise-optimization-for-nonlinear-recurrent-memory.