Authors

Kaan Ozkara

Publications 1

Preprint Open access

Optimizing Large Language Models with Chained LMOs

Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chai …

Co-authors