TensorX
返回文献探索

Paper · arXiv 2407.10960

Fast Matrix Multiplications for Lookup Table-Quantized LLMs

Han Guo, William Brandon, Radostin Cholakov, Jonathan Ragan-Kelley, Eric P. Xing, Yoon Kim

13 upvotesJuly 15, 2024arXiv 预印本
AI 摘要

FLUTE, a lookup table engine for quantized language models, accelerates inference by minimizing bit manipulations and optimizing shared memory usage, achieving faster performance compared to existing kernels.

large language modelsweight-only quantizationlookup tablequantized weight matrixvectorizationGEMM kernelsNormalFloat quantizationLLaMA3

Abstract

The deployment of large language models (LLMs) is often constrained by memory bandwidth, where the primary bottleneck is the cost of transferring model parameters from the GPU's global memory to its registers. When coupled with custom kernels that fuse the dequantization and matmul operations, weight-only quantization can thus enable faster inference by reducing the amount of memory movement. However, developing high-performance kernels for weight-quantized LLMs presents substantial challenges, especially when the weights are compressed to non-evenly-divisible bit widths (e.g., 3 bits) with non-uniform, lookup table (LUT) quantization. This paper describes FLUTE, a flexible lookup table engine for LUT-quantized LLMs, which uses offline restructuring of the quantized weight matrix to minimize bit manipulations associated with unpacking, and vectorization and duplication of the lookup table to mitigate shared memory bandwidth constraints. At batch sizes < 32 and quantization group size of 128 (typical in LLM inference), the FLUTE kernel can be 2-4x faster than existing GEMM kernels. As an application of FLUTE, we explore a simple extension to lookup table-based NormalFloat quantization and apply it to quantize LLaMA3 to various configurations, obtaining competitive quantization performance against strong baselines while obtaining an end-to-end throughput increase of 1.5 to 2 times.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Fast Matrix Multiplications for Lookup Table-Quantized LLMs | TensorX