الباحثون

Jeff Gore

المنشورات 3

نسخة أولية وصول مفتوح

Emergent Inverse-Depth Scaling From Nonlinearity In Attention

Zirui Peng, Yizhou Liu, Ziming Liu وآخرون · 2026

Scaling laws describe power-law improvements in model performance with dataset size and parameter count, yet their underlying mechanisms are not fully understood. To explain the parameter count scaling, existing theory posits power-law scaling with model depth. In linear-attention models, this scaling is tied to a powe …

نسخة أولية وصول مفتوح

Optimizer-dependent training dynamics converge to the same one-third optimal data scaling

Hyunseok Lee, Mihir Basil, Yizhou Liu وآخرون · 2026

Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the …

المؤلفون المشاركون