Abstract
Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size $D$. We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to $1/3$ across optimizers. In an online teacher-student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents $α_{r}$ and $α_{t}$. Under SGD, both are close to $1/3$, so the data exponent is also $1/3$ across different learning rates. Under Adam the two separate: $α_{r} \simeq 0.48$ but $α_{t} \simeq 0.08$. Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: $D$-independent for SGD but falls with $D$ for Adam. Yet tuned to that optimum, the loss returns to $D^{-1/3}$ for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, $2α_{r}+ α_{t} = 1$, which fixes the optimal data exponent at $1/3$. Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with $D^{-1/3}$ across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Lee, H., Basil, M., Liu, Y., & Gore, J. (2026). Optimizer-dependent training dynamics converge to the same one-third optimal data scaling. https://omanscience.com/en/articles/optimizer-dependent-training-dynamics-converge-to-the-same-one-third-optimal-data-scaling
MLA 9
Lee, Hyunseok, et al. "Optimizer-dependent training dynamics converge to the same one-third optimal data scaling." https://omanscience.com/en/articles/optimizer-dependent-training-dynamics-converge-to-the-same-one-third-optimal-data-scaling.
Chicago (author–date)
Lee, Hyunseok, Mihir Basil, Yizhou Liu, and Jeff Gore. 2026. "Optimizer-dependent training dynamics converge to the same one-third optimal data scaling." https://omanscience.com/en/articles/optimizer-dependent-training-dynamics-converge-to-the-same-one-third-optimal-data-scaling.
Harvard
Lee, H., Basil, M., Liu, Y. and Gore, J. (2026) 'Optimizer-dependent training dynamics converge to the same one-third optimal data scaling', Available at: https://omanscience.com/en/articles/optimizer-dependent-training-dynamics-converge-to-the-same-one-third-optimal-data-scaling.
Vancouver
Lee H, Basil M, Liu Y, Gore J. Optimizer-dependent training dynamics converge to the same one-third optimal data scaling. https://omanscience.com/en/articles/optimizer-dependent-training-dynamics-converge-to-the-same-one-third-optimal-data-scaling
IEEE
H. Lee, M. Basil, Y. Liu, and J. Gore, "Optimizer-dependent training dynamics converge to the same one-third optimal data scaling," https://omanscience.com/en/articles/optimizer-dependent-training-dynamics-converge-to-the-same-one-third-optimal-data-scaling.