TensorX
返回文献探索

Paper · arXiv 2507.08799

KV Cache Steering for Inducing Reasoning in Small Language Models

Max Belitsky, Dawid J. Kopiczko, Michael Dorkenwald, M. Jehanzeb Mirza, Cees G. M. Snoek, Yuki M. Asano

40 upvotesJuly 11, 2025arXiv 预印本
AI 摘要

Cache steering improves language model reasoning through a single intervention in the key-value cache, enhancing both structure and performance without fine-tuning.

cache steeringkey-value cachechain-of-thought reasoningGPT-4steering vectorsmulti-step reasoningactivation steeringhyperparameter stabilityinference-time efficiencyease of integrationcontrolled generation

Abstract

We propose cache steering, a lightweight method for implicit steering of language models via a one-shot intervention applied directly to the key-value cache. To validate its effectiveness, we apply cache steering to induce chain-of-thought reasoning in small language models. Our approach leverages GPT-4o-generated reasoning traces to construct steering vectors that shift model behavior toward more explicit, multi-step reasoning without fine-tuning or prompt modifications. Experimental evaluations on diverse reasoning benchmarks demonstrate that cache steering improves both the qualitative structure of model reasoning and quantitative task performance. Compared to prior activation steering techniques that require continuous interventions, our one-shot cache steering offers substantial advantages in terms of hyperparameter stability, inference-time efficiency, and ease of integration, making it a more robust and practical solution for controlled generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
KV Cache Steering for Inducing Reasoning in Small Language Models | TensorX