TensorX
返回文献探索

Paper · arXiv 2406.11430

A Simple and Effective L_2 Norm-Based Strategy for KV Cache Compression

Alessio Devoto, Yu Zhao, Simone Scardapane, Pasquale Minervini

25 upvotesJune 17, 2024arXiv 预印本
AI 摘要

By compressing the KV cache based on the $L_2$ norm of key embeddings, the sizes required for transformer models can be significantly reduced without sacrificing accuracy.

large language modelskey-value cachedecoder-only Transformersattention distributions$L_2$ normkey embeddinglanguage modellingneedle-in-a-haystack taskspasskey retrieval tasks

Abstract

The deployment of large language models (LLMs) is often hindered by the extensive memory requirements of the Key-Value (KV) cache, especially as context lengths increase. Existing approaches to reduce the KV cache size involve either fine-tuning the model to learn a compression strategy or leveraging attention scores to reduce the sequence length. We analyse the attention distributions in decoder-only Transformers-based models and observe that attention allocation patterns stay consistent across most layers. Surprisingly, we find a clear correlation between the L_2 and the attention scores over cached KV pairs, where a low L_2 of a key embedding usually leads to a high attention score during decoding. This finding indicates that the influence of a KV pair is potentially determined by the key embedding itself before being queried. Based on this observation, we compress the KV cache based on the L_2 of key embeddings. Our experimental results show that this simple strategy can reduce the KV cache size by 50% on language modelling and needle-in-a-haystack tasks and 90% on passkey retrieval tasks without losing accuracy.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号