TensorX
返回文献探索

Paper · arXiv 2501.16937

TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models

Makoto Shing, Kou Misaki, Han Bao, Sho Yokoi, Takuya Akiba

8 upvotesJanuary 28, 2025arXiv 预印本
AI 摘要

Temporally Adaptive Interpolated Distillation (TAID) addresses capacity gap, mode averaging, and mode collapse between large and small models, enhancing knowledge distillation and leading to efficient compact foundation models for language and vision-language tasks.

causal language modelsknowledge distillationteacher modelstudent modelcapacity gapmode averagingmode collapseTemporally Adaptive Interpolated Distillation (TAID)adaptive intermediate distributioninstruction tuningpre-trainingTAID-LLM-1.5BTAID-VLM-2B

Abstract

Causal language models have demonstrated remarkable capabilities, but their size poses significant challenges for deployment in resource-constrained environments. Knowledge distillation, a widely-used technique for transferring knowledge from a large teacher model to a small student model, presents a promising approach for model compression. A significant remaining issue lies in the major differences between teacher and student models, namely the substantial capacity gap, mode averaging, and mode collapse, which pose barriers during distillation. To address these issues, we introduce Temporally Adaptive Interpolated Distillation (TAID), a novel knowledge distillation approach that dynamically interpolates student and teacher distributions through an adaptive intermediate distribution, gradually shifting from the student's initial distribution towards the teacher's distribution. We provide a theoretical analysis demonstrating TAID's ability to prevent mode collapse and empirically show its effectiveness in addressing the capacity gap while balancing mode averaging and mode collapse. Our comprehensive experiments demonstrate TAID's superior performance across various model sizes and architectures in both instruction tuning and pre-training scenarios. Furthermore, we showcase TAID's practical impact by developing two state-of-the-art compact foundation models: TAID-LLM-1.5B for language tasks and TAID-VLM-2B for vision-language tasks. These results demonstrate TAID's effectiveness in creating high-performing and efficient models, advancing the development of more accessible AI technologies.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号