TensorX
返回文献探索

Paper · arXiv 2307.09638

Promoting Exploration in Memory-Augmented Adam using Critical Momenta

Pranshu Malviya, Gonçalo Mordido, Aristide Baratin, Reza Babanezhad Harikandeh, Jerry Huang, Simon Lacoste-Julien, Razvan Pascanu, Sarath Chandar

2 upvotesJuly 18, 2023arXiv 预印本
AI 摘要

A memory-augmented version of Adam optimizer enhances generalization by encouraging exploration towards flatter minima in the loss landscape.

gradient-based optimizersAdamfast convergencehyperparameter choiceflat minimasharp basinsloss landscapemomentum termsbufferexplorationoverhootingsupervised language modellingimage classification

Abstract

Adaptive gradient-based optimizers, particularly Adam, have left their mark in training large-scale deep learning models. The strength of such optimizers is that they exhibit fast convergence while being more robust to hyperparameter choice. However, they often generalize worse than non-adaptive methods. Recent studies have tied this performance gap to flat minima selection: adaptive methods tend to find solutions in sharper basins of the loss landscape, which in turn hurts generalization. To overcome this issue, we propose a new memory-augmented version of Adam that promotes exploration towards flatter minima by using a buffer of critical momentum terms during training. Intuitively, the use of the buffer makes the optimizer overshoot outside the basin of attraction if it is not wide enough. We empirically show that our method improves the performance of several variants of Adam on standard supervised language modelling and image classification tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Promoting Exploration in Memory-Augmented Adam using Critical Momenta | TensorX