TensorX
返回文献探索

Paper · arXiv 2307.05695

Stack More Layers Differently: High-Rank Training Through Low-Rank Updates

Vladislav Lialin, Namrata Shivagunde, Sherin Muckatira, Anna Rumshisky

25 upvotesJuly 11, 2023arXiv 预印本
AI 摘要

ReLoRA, a low-rank training method, demonstrates comparable performance to traditional training for large transformer language models and becomes more efficient with increased model size.

low-rank trainingReLoRAhigh-rank networkstransformer language modelsscaling laws

Abstract

Despite the dominance and effectiveness of scaling, resulting in large networks with hundreds of billions of parameters, the necessity to train overparametrized models remains poorly understood, and alternative approaches do not necessarily make it cheaper to train high-performance models. In this paper, we explore low-rank training techniques as an alternative approach to training large neural networks. We introduce a novel method called ReLoRA, which utilizes low-rank updates to train high-rank networks. We apply ReLoRA to pre-training transformer language models with up to 350M parameters and demonstrate comparable performance to regular neural network training. Furthermore, we observe that the efficiency of ReLoRA increases with model size, making it a promising approach for training multi-billion-parameter networks efficiently. Our findings shed light on the potential of low-rank training techniques and their implications for scaling laws.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Stack More Layers Differently: High-Rank Training Through Low-Rank Updates | TensorX