TensorX
返回文献探索

Paper · arXiv 2502.16894

Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment

Chenghao Fan, Zhenyi Lu, Sichen Liu, Xiaoye Qu, Wei Wei, Chengfeng Gu, Yu Cheng

33 upvotesFebruary 24, 2025arXiv 预印本
AI 摘要

GOAT, a framework that integrates relevant priors using an SVD-structured Mixture-of-Experts and derives a theoretical scaling factor, enhances LoRA MoE's efficiency and performance, approaching Full Fine-Tuning results.

Low-Rank AdaptationLoRAParameter-efficient fine-tuningLarge Language ModelsFull Fine-TuningSingular value decompositionSVDMixture-of-ExpertsMoEWeight misalignmentGradient dynamicsSVD-structured MoETheoretical scaling factor

Abstract

While Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning for Large Language Models (LLMs), its performance often falls short of Full Fine-Tuning (Full FT). Current methods optimize LoRA by initializing with static singular value decomposition (SVD) subsets, leading to suboptimal leveraging of pre-trained knowledge. Another path for improving LoRA is incorporating a Mixture-of-Experts (MoE) architecture. However, weight misalignment and complex gradient dynamics make it challenging to adopt SVD prior to the LoRA MoE architecture. To mitigate these issues, we propose Great LoRA Mixture-of-Expert (GOAT), a framework that (1) adaptively integrates relevant priors using an SVD-structured MoE, and (2) aligns optimization with full fine-tuned MoE by deriving a theoretical scaling factor. We demonstrate that proper scaling, without modifying the architecture or training algorithms, boosts LoRA MoE's efficiency and performance. Experiments across 25 datasets, including natural language understanding, commonsense reasoning, image classification, and natural language generation, demonstrate GOAT's state-of-the-art performance, closing the gap with Full FT.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号