TensorX
返回文献探索

Paper · arXiv 2410.16215

Pre-training Distillation for Large Language Models: A Design Space Exploration

Hao Peng, Xin Lv, Yushi Bai, Zijun Yao, Jiajie Zhang, Lei Hou, Juanzi Li

18 upvotesOctober 21, 2024arXiv 预印本
AI 摘要

The study extends knowledge distillation to the pre-training phase of large language models, demonstrating its effectiveness and exploring key design factors like logits processing, loss selection, scaling laws, and offline/online logits.

knowledge distillationlarge language modelspre-training distillationGLM-4-9Blogits processingloss selectionscaling lawoffline or online logits

Abstract

Knowledge distillation (KD) aims to transfer knowledge from a large teacher model to a smaller student model. Previous work applying KD in the field of large language models (LLMs) typically focused on the post-training phase, where the student LLM learns directly from instructions and corresponding responses generated by the teacher model. In this paper, we extend KD to the pre-training phase of LLMs, named pre-training distillation (PD). We first conduct a preliminary experiment using GLM-4-9B as the teacher LLM to distill a 1.9B parameter student LLM, validating the effectiveness of PD. Considering the key impact factors of distillation, we systematically explore the design space of pre-training distillation across four aspects: logits processing, loss selection, scaling law, and offline or online logits. We conduct extensive experiments to explore the design space of pre-training distillation and find better configurations and interesting conclusions, such as larger student LLMs generally benefiting more from pre-training distillation, while a larger teacher LLM does not necessarily guarantee better results. We hope our exploration of the design space will inform future practices in pre-training distillation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Pre-training Distillation for Large Language Models: A Design Space Exploration | TensorX