TensorX
返回文献探索

Paper · arXiv 2310.08659

LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models

Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, Tuo Zhao

30 upvotesOctober 12, 2023arXiv 预印本
AI 摘要

LoftQ, a novel quantization framework, improves performance of LLMs by simultaneously quantizing and initializing LoRA fine-tuning, reducing the gap in downstream task performance.

quantizationLarge Language Models (LLMs)LoRA fine-tuningLoftQlow-rank initializationgeneralizationnatural language understandingquestion answeringsummarizationnatural language generation2-bit2/4-bit mixed precision

Abstract

Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning. In this work we focus on the scenario where quantization and LoRA fine-tuning are applied together on a pre-trained model. In such cases it is common to observe a consistent gap in the performance on downstream tasks between full fine-tuning and quantization plus LoRA fine-tuning approach. In response, we propose LoftQ (LoRA-Fine-Tuning-aware Quantization), a novel quantization framework that simultaneously quantizes an LLM and finds a proper low-rank initialization for LoRA fine-tuning. Such an initialization alleviates the discrepancy between the quantized and full-precision model and significantly improves the generalization in downstream tasks. We evaluate our method on natural language understanding, question answering, summarization, and natural language generation tasks. Experiments show that our method is highly effective and outperforms existing quantization methods, especially in the challenging 2-bit and 2/4-bit mixed precision regimes. We will release our code.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号