Abstract

Large learning rates can qualitatively change the trajectory of neural network training, often pushing optimization into regimes far from classical gradient-flow behavior. The Edge of Stability (EoS) offers a valuable lens on the dynamics such learning rates induce. We study corresponding dynamics in diagonal linear networks, where we uncover a competition between two distinct implicit biases that jointly determine the sparsity of the recovered solution in regression settings. Complementary to the Gain, which captures the average discretization error accumulated by Gradient Descent relative to Gradient Flow, we derive a closely associated but overlooked quantity: the Drift. Under large learning rates, it describes an imbalance between different discretization errors and represents a systematic shift in the optimization trajectory. While the Gain grows monotonically in certain regimes, and can bias towards denser, flatter interpolators, the impact of the Drift depends on its alignment with potential solutions, which can either counteract or reinforce the effect of the Gain. Consequently, its behavior drives model selection, particularly during early training epochs. To validate our theoretical insights, we introduce an intervention that actively steers the Gain to recover sharper, sparser solutions. Thus, our analysis reveals that large learning rates do not universally hinder the recovery of sparse solutions. On the contrary, they can be harnessed to control the implicit bias of training.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Sanyal, A., Jacobs, T., & Burkholz, R. (2026). Mind the Drift: Diagonal Linear Networks Under Large Learning Rates. https://omanscience.com/en/articles/mind-the-drift-diagonal-linear-networks-under-large-learning-rates

MLA 9

Sanyal, Aniket, et al. "Mind the Drift: Diagonal Linear Networks Under Large Learning Rates." https://omanscience.com/en/articles/mind-the-drift-diagonal-linear-networks-under-large-learning-rates.

Chicago (author–date)

Sanyal, Aniket, Tom Jacobs, and Rebekka Burkholz. 2026. "Mind the Drift: Diagonal Linear Networks Under Large Learning Rates." https://omanscience.com/en/articles/mind-the-drift-diagonal-linear-networks-under-large-learning-rates.

Harvard

Sanyal, A., Jacobs, T. and Burkholz, R. (2026) 'Mind the Drift: Diagonal Linear Networks Under Large Learning Rates', Available at: https://omanscience.com/en/articles/mind-the-drift-diagonal-linear-networks-under-large-learning-rates.

Vancouver

Sanyal A, Jacobs T, Burkholz R. Mind the Drift: Diagonal Linear Networks Under Large Learning Rates. https://omanscience.com/en/articles/mind-the-drift-diagonal-linear-networks-under-large-learning-rates

IEEE

A. Sanyal, T. Jacobs, and R. Burkholz, "Mind the Drift: Diagonal Linear Networks Under Large Learning Rates," https://omanscience.com/en/articles/mind-the-drift-diagonal-linear-networks-under-large-learning-rates.