الملخص
Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Kim, S., Ozkara, K., & Park, Y. (2026). Optimizing Large Language Models with Chained LMOs. https://omanscience.com/ar/articles/optimizing-large-language-models-with-chained-lmos
MLA 9
Kim, Sungyoon, et al. "Optimizing Large Language Models with Chained LMOs." https://omanscience.com/ar/articles/optimizing-large-language-models-with-chained-lmos.
شيكاغو (المؤلف–التاريخ)
Kim, Sungyoon, Kaan Ozkara, and Youngsuk Park. 2026. "Optimizing Large Language Models with Chained LMOs." https://omanscience.com/ar/articles/optimizing-large-language-models-with-chained-lmos.
هارفارد
Kim, S., Ozkara, K. and Park, Y. (2026) 'Optimizing Large Language Models with Chained LMOs', Available at: https://omanscience.com/ar/articles/optimizing-large-language-models-with-chained-lmos.
فانكوفر
Kim S, Ozkara K, Park Y. Optimizing Large Language Models with Chained LMOs. https://omanscience.com/ar/articles/optimizing-large-language-models-with-chained-lmos
IEEE
S. Kim, K. Ozkara, and Y. Park, "Optimizing Large Language Models with Chained LMOs," https://omanscience.com/ar/articles/optimizing-large-language-models-with-chained-lmos.