TensorX
返回文献探索

Paper · arXiv 2502.08606

Distillation Scaling Laws

Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb

48 upvotesFebruary 12, 2025arXiv 预印本
AI 摘要

The study presents a scaling law to optimize compute allocation for model distillation, showing conditions under which distillation outperforms supervised pretraining.

distillation scaling lawdistilled modelcompute budgetstudent modelteacher modelcompute optimal distillationsupervised pretraininglarge scale study

Abstract

We provide a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings reduce the risks associated with using distillation at scale; compute allocation for both the teacher and student models can now be done to maximize student performance. We provide compute optimal distillation recipes for when 1) a teacher exists, or 2) a teacher needs training. If many students are to be distilled, or a teacher already exists, distillation outperforms supervised pretraining until a compute level which grows predictably with student size. If one student is to be distilled and a teacher also needs training, supervised learning should be done instead. Additionally, we provide insights across our large scale study of distillation, which increase our understanding of distillation and inform experimental design.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号