TensorX
返回文献探索

Paper · arXiv 2407.01906

Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models

Zihan Wang, Deli Chen, Damai Dai, Runxin Xu, Zhuoshu Li, Y. Wu

47 upvotesJuly 2, 2024arXiv 预印本
AI 摘要

Research on parameter-efficient fine-tuning for Large Language Models with Mixture-of-Experts architecture reveals that task-specific expert tuning can enhance performance and efficiency.

Parameter-efficient fine-tuningPEFTLarge Language ModelsLLMsMixture-of-ExpertsMoEdispersion degreeactivated expertsrouting distributionExpert-Specialized Fine-TuningESFTfull-parameter fine-tuningtraining efficiencyeffectiveness

Abstract

Parameter-efficient fine-tuning (PEFT) is crucial for customizing Large Language Models (LLMs) with constrained resources. Although there have been various PEFT methods for dense-architecture LLMs, PEFT for sparse-architecture LLMs is still underexplored. In this work, we study the PEFT method for LLMs with the Mixture-of-Experts (MoE) architecture and the contents of this work are mainly threefold: (1) We investigate the dispersion degree of the activated experts in customized tasks, and found that the routing distribution for a specific task tends to be highly concentrated, while the distribution of activated experts varies significantly across different tasks. (2) We propose Expert-Specialized Fine-Tuning, or ESFT, which tunes the experts most relevant to downstream tasks while freezing the other experts and modules; experimental results demonstrate that our method not only improves the tuning efficiency, but also matches or even surpasses the performance of full-parameter fine-tuning. (3) We further analyze the impact of the MoE architecture on expert-specialized fine-tuning. We find that MoE models with finer-grained experts are more advantageous in selecting the combination of experts that are most relevant to downstream tasks, thereby enhancing both the training efficiency and effectiveness.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号