TensorX
返回文献探索

Paper · arXiv 2607.05061

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl, David Stap, Thomas Schmied, Sebastian Böck, Günter Klambauer, Sepp Hochreiter

24 upvotesJuly 6, 2026arXiv 预印本
AI 摘要

KVpop learns optimal key-value cache eviction by directly supervising keep-or-drop decisions using future-attention targets, achieving high performance with reduced memory usage.

KV cacheautoregressive decodingKV evictionfuture-attention targetdelayed memory-based scorerattention mapsKV cache compressionQwen3-4BQwen3-8B

Abstract

Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, which poorly track future token utility and cause brittle eviction as relevance shifts. To address this, we introduce KVpop, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision. The scorer is trained against a novel future-attention target, computed efficiently without materializing dense attention maps. We further introduce a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context. On AIME and HMMT mathematical reasoning, KVpop retains 98% of full-attention performance on Qwen3-4B at 75% KV cache compression and 97% at 88% compression, consistently outperforming established eviction baselines. Qwen3-8B shows even stronger results, reaching near-full teacher performance. These results show that supervising eviction with future-attention signals cuts memory costs while maintaining quality.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
KVpop -- Key-Value Cache Compression with Predictive Online Pruning | TensorX