TensorX
返回文献探索

Paper · arXiv 2312.07046

Rethinking Compression: Reduced Order Modelling of Latent Features in Large Language Models

Arnav Chavan, Nahush Lele, Deepak Gupta

13 upvotesDecember 12, 2023arXiv 预印本
AI 摘要

This paper presents a new compression technique for Large Language Models (LLMs) using reduced order modeling and low-rank decomposition that operates in a layer-wise manner without a GPU, achieving better results than structured pruning.

Large Language Models (LLMs)reduced order modelinglow-rank decompositionweight spacelayer-wise compressionmatrix decompositionstructured pruning

Abstract

Due to the substantial scale of Large Language Models (LLMs), the direct application of conventional compression methodologies proves impractical. The computational demands associated with even minimal gradient updates present challenges, particularly on consumer-grade hardware. This paper introduces an innovative approach for the parametric and practical compression of LLMs based on reduced order modelling, which entails low-rank decomposition within the feature space and re-parameterization in the weight space. Notably, this compression technique operates in a layer-wise manner, obviating the need for a GPU device and enabling the compression of billion-scale models within stringent constraints of both memory and time. Our method represents a significant advancement in model compression by leveraging matrix decomposition, demonstrating superior efficacy compared to the prevailing state-of-the-art structured pruning method.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Rethinking Compression: Reduced Order Modelling of Latent Features in Large Language Models | TensorX