TensorX
返回文献探索

Paper · arXiv 2305.02790

BranchNorm: Robustly Scaling Extremely Deep Transformers

Yijin Liu, Xianfeng Zeng, Fandong Meng, Jie Zhou

1 upvotesMay 4, 2023arXiv 预印本
AI 摘要

BranchNorm dynamically adjusts the scaling of the non-residual branch in Transformers to improve training stability and convergence in deep models.

DeepNormTransformersdeep scalingmodel updateBranchNormgradient normstraining stabilityconvergence

Abstract

Recently, DeepNorm scales Transformers into extremely deep (i.e., 1000 layers) and reveals the promising potential of deep scaling. To stabilize the training of deep models, DeepNorm (Wang et al., 2022) attempts to constrain the model update to a constant value. Although applying such a constraint can benefit the early stage of model training, it may lead to undertrained models during the whole training procedure. In this paper, we propose BranchNorm, which dynamically rescales the non-residual branch of Transformer in accordance with the training period. BranchNorm not only theoretically stabilizes the training with smooth gradient norms at the early stage, but also encourages better convergence in the subsequent training stage. Experiment results on multiple translation tasks demonstrate that BranchNorm achieves a better trade-off between training stability and converge performance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
BranchNorm: Robustly Scaling Extremely Deep Transformers | TensorX