TensorX
返回文献探索

Paper · arXiv 2403.05812

Algorithmic progress in language models

Anson Ho, Tamay Besiroglu, Ege Erdil, David Owen, Robi Rahman, Zifan Carl Guo, David Atkinson, Neil Thompson, Jaime Sevilla

19 upvotesMarch 9, 2024arXiv 预印本
AI 摘要

Analysis of language model evaluations shows that compute requirements halve faster than hardware improvements, with compute making the larger contribution to performance gains.

pre-training language modelsdeep learningdatasetWikitextPenn Treebankcomputeperformance thresholdMoore's Lawaugmented scaling lawstraining algorithmstransformerlanguage modeling

Abstract

We investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning. Using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning 2012-2023, we find that the compute required to reach a set performance threshold has halved approximately every 8 months, with a 95% confidence interval of around 5 to 14 months, substantially faster than hardware gains per Moore's Law. We estimate augmented scaling laws, which enable us to quantify algorithmic progress and determine the relative contributions of scaling models versus innovations in training algorithms. Despite the rapid pace of algorithmic progress and the development of new architectures such as the transformer, our analysis reveals that the increase in compute made an even larger contribution to overall performance improvements over this time period. Though limited by noisy benchmark data, our analysis quantifies the rapid progress in language modeling, shedding light on the relative contributions from compute and algorithms.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号