TensorX
返回文献探索

Paper · arXiv 2405.15682

The Road Less Scheduled

Aaron Defazio, Xingyu, Yang, Harsh Mehta, Konstantin Mishchenko, Ahmed Khaled, Ashok Cutkosky

27 upvotesMay 24, 2024arXiv 预印本
AI 摘要

A new schedule-free method achieves state-of-the-art performance in optimization without requiring a stopping time or additional hyper-parameters, by unifying scheduling and iterate averaging.

learning rate schedulesoptimization stopping steplearning ratestate-of-the-arthyper-parametersmomentumiterate averagingschedule-free

Abstract

Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We propose an approach that avoids the need for this stopping time by eschewing the use of schedules entirely, while exhibiting state-of-the-art performance compared to schedules across a wide family of problems ranging from convex problems to large-scale deep learning problems. Our Schedule-Free approach introduces no additional hyper-parameters over standard optimizers with momentum. Our method is a direct consequence of a new theory we develop that unifies scheduling and iterate averaging. An open source implementation of our method is available (https://github.com/facebookresearch/schedule_free).

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
The Road Less Scheduled | TensorX