TensorX
返回文献探索

Paper · arXiv 2501.16372

Low-Rank Adapters Meet Neural Architecture Search for LLM Compression

J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain

11 upvotesJanuary 23, 2025arXiv 预印本
AI 摘要

Low-rank adapters and NAS techniques are combined to create efficient, parameter-reduced Large Language Models suitable for resource-limited environments.

Large Language Models (LLMs)low-rank adaptersparameter-efficient fine-tuning (PEFT)Neural Architecture Search (NAS)weight-sharing super-networksmemory footprintsinference times

Abstract

The rapid expansion of Large Language Models (LLMs) has posed significant challenges regarding the computational resources required for fine-tuning and deployment. Recent advancements in low-rank adapters have demonstrated their efficacy in parameter-efficient fine-tuning (PEFT) of these models. This retrospective paper comprehensively discusses innovative approaches that synergize low-rank representations with Neural Architecture Search (NAS) techniques, particularly weight-sharing super-networks. Robust solutions for compressing and fine-tuning large pre-trained models are developed by integrating these methodologies. Our analysis highlights the potential of these combined strategies to democratize the use of LLMs, making them more accessible for deployment in resource-constrained environments. The resulting models exhibit reduced memory footprints and faster inference times, paving the way for more practical and scalable applications of LLMs. Models and code are available at https://github.com/IntelLabs/Hardware-Aware-Automated-Machine-Learning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号