TensorX
返回文献探索

Paper · arXiv 2502.03032

Analyze Feature Flow to Enhance Interpretation and Steering in Language Models

Daniil Laptev, Nikita Balagansky, Yaroslav Aksenov, Daniil Gavrilov

59 upvotesFebruary 5, 2025arXiv 预印本
AI 摘要

A new data-free cosine similarity method maps and steers features across layers of large language models, enhancing interpretability and targeted control in text generation.

sparse autoencoderlarge language modelsdata-free cosine similarityfeature evolutionforward passes

Abstract

We introduce a new approach to systematically map features discovered by sparse autoencoder across consecutive layers of large language models, extending earlier work that examined inter-layer feature links. By using a data-free cosine similarity technique, we trace how specific features persist, transform, or first appear at each stage. This method yields granular flow graphs of feature evolution, enabling fine-grained interpretability and mechanistic insights into model computations. Crucially, we demonstrate how these cross-layer feature maps facilitate direct steering of model behavior by amplifying or suppressing chosen features, achieving targeted thematic control in text generation. Together, our findings highlight the utility of a causal, cross-layer interpretability framework that not only clarifies how features develop through forward passes but also provides new means for transparent manipulation of large language models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models | TensorX