TensorX
返回文献探索

Paper · arXiv 2401.00448

Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Nikhil Sardana, Jonathan Frankle

31 upvotesDecember 31, 2023arXiv 预印本
AI 摘要

The modified Chinchilla scaling laws incorporate inference cost, suggesting that models for large inference demand should be smaller and trained longer than originally proposed.

Large language modelscaling lawsDeepMind Chinchillaparameter countpre-training datainference demandcompute budgetreal-world costsChinchilla-optimal

Abstract

Large language model (LLM) scaling laws are empirical formulas that estimate changes in model quality as a result of increasing parameter count and training data. However, these formulas, including the popular DeepMind Chinchilla scaling laws, neglect to include the cost of inference. We modify the Chinchilla scaling laws to calculate the optimal LLM parameter count and pre-training data size to train and deploy a model of a given quality and inference demand. We conduct our analysis both in terms of a compute budget and real-world costs and find that LLM researchers expecting reasonably large inference demand (~1B requests) should train models smaller and longer than Chinchilla-optimal.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws | TensorX