TensorX
返回文献探索

Paper · arXiv 2505.02222

Practical Efficiency of Muon for Pretraining

Essential AI, Ishaan Shah, Anthony M. Polloreno, Karl Stratos, Philip Monk, Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, Anil Thomas, Ashish Tanwer, Darsh J Shah, Khoi Nguyen, Kurt Smith, Michael Callahan, Michael Pust, Mohit Parmar, Peter Rushton, Platon Mazarakis, Ritvik Kapila, Saurabh Srivastava, Somanshu Singla, Tim Romanski, Yash Vanjani, Ashish Vaswani

42 upvotesMay 4, 2025arXiv 预印本
AI 摘要

Muon, a second-order optimizer, improves data efficiency and computational savings over AdamW, especially at large batch sizes, and combined with muP, it provides efficient hyperparameter transfer and minimal resource overhead.

Muonsecond-order optimizerPareto frontierAdamWdata efficiencycritical batch sizemaximal update parameterizationmuPhyperparameter transfertelescoping algorithmparametersdata distributionarchitecture

Abstract

We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号