الملخص
A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Dang, X., Wen, K., & Malladi, S. (2026). The Best Optimizer Depends on Batch Size. https://omanscience.com/ar/articles/the-best-optimizer-depends-on-batch-size
MLA 9
Dang, Xingyu, et al. "The Best Optimizer Depends on Batch Size." https://omanscience.com/ar/articles/the-best-optimizer-depends-on-batch-size.
شيكاغو (المؤلف–التاريخ)
Dang, Xingyu, Kaiyue Wen, and Sadhika Malladi. 2026. "The Best Optimizer Depends on Batch Size." https://omanscience.com/ar/articles/the-best-optimizer-depends-on-batch-size.
هارفارد
Dang, X., Wen, K. and Malladi, S. (2026) 'The Best Optimizer Depends on Batch Size', Available at: https://omanscience.com/ar/articles/the-best-optimizer-depends-on-batch-size.
فانكوفر
Dang X, Wen K, Malladi S. The Best Optimizer Depends on Batch Size. https://omanscience.com/ar/articles/the-best-optimizer-depends-on-batch-size
IEEE
X. Dang, K. Wen, and S. Malladi, "The Best Optimizer Depends on Batch Size," https://omanscience.com/ar/articles/the-best-optimizer-depends-on-batch-size.