TensorX
返回文献探索

Paper · arXiv 2406.00153

μLO: Compute-Efficient Meta-Generalization of Learned Optimizers

Benjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon, Irina Rish, Eugene Belilovsky

12 upvotesMay 31, 2024arXiv 预印本
AI 摘要

Maximal Update Parametrization improves the meta-generalization of learned optimizers, allowing zero-shot generalization to larger and deeper models with reduced computational cost.

Learned optimizersMaximal Update Parametrizationmeta-generalizationzero-shot generalizationlarge-width modelsdeeper networkstraining horizons

Abstract

Learned optimizers (LOs) can significantly reduce the wall-clock training time of neural networks, substantially reducing training costs. However, they often suffer from poor meta-generalization, especially when training networks larger than those seen during meta-training. To address this, we use the recently proposed Maximal Update Parametrization (muP), which allows zero-shot generalization of optimizer hyperparameters from smaller to larger models. We extend muP theory to learned optimizers, treating the meta-training problem as finding the learned optimizer under muP. Our evaluation shows that LOs meta-trained with muP substantially improve meta-generalization as compared to LOs trained under standard parametrization (SP). Notably, when applied to large-width models, our best muLO, trained for 103 GPU-hours, matches or exceeds the performance of VeLO, the largest publicly available learned optimizer, meta-trained with 4000 TPU-months of compute. Moreover, muLOs demonstrate better generalization than their SP counterparts to deeper networks and to much longer training horizons (25 times longer) than those seen during meta-training.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
μLO: Compute-Efficient Meta-Generalization of Learned Optimizers | TensorX