TensorX
返回文献探索

Paper · arXiv 2309.08520

Scaling Laws for Sparsely-Connected Foundation Models

Elias Frantar, Carlos Riquelme, Neil Houlsby, Dan Alistarh, Utku Evci

15 upvotesSeptember 15, 2023arXiv 预印本
AI 摘要

Parameter sparsity in Transformers affects their scaling behavior on large datasets; optimal sparsity increases with data volume and different structures can impact performance.

parameter sparsityscaling lawTransformersweight sparsityViTJFT-4BT5C4optimal sparsityhardware-friendly n:m patternpretrained dense model

Abstract

We explore the impact of parameter sparsity on the scaling behavior of Transformers trained on massive datasets (i.e., "foundation models"), in both vision and language domains. In this setting, we identify the first scaling law describing the relationship between weight sparsity, number of non-zero parameters, and amount of training data, which we validate empirically across model and data scales; on ViT/JFT-4B and T5/C4. These results allow us to characterize the "optimal sparsity", the sparsity level which yields the best performance for a given effective model size and training budget. For a fixed number of non-zero parameters, we identify that the optimal sparsity increases with the amount of data used for training. We also extend our study to different sparsity structures (such as the hardware-friendly n:m pattern) and strategies (such as starting from a pretrained dense model). Our findings shed light on the power and limitations of weight sparsity across various parameter and computational settings, offering both theoretical understanding and practical implications for leveraging sparsity towards computational efficiency improvements.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Scaling Laws for Sparsely-Connected Foundation Models | TensorX