TensorX
返回文献探索

Paper · arXiv 2309.08600

Sparse Autoencoders Find Highly Interpretable Features in Language Models

Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, Lee Sharkey

15 upvotesSeptember 15, 2023arXiv 预印本
AI 摘要

Using sparse autoencoders, this work identifies and removes overcomplete directions in language model activations, improving interpretability and model editability.

polysemanticitysuperpositionsparse autoencodersinternal activationsmonosemanticinterpretabilitymodel editingpronoun prediction

Abstract

One of the roadblocks to a better understanding of neural networks' internals is polysemanticity, where neurons appear to activate in multiple, semantically distinct contexts. Polysemanticity prevents us from identifying concise, human-understandable explanations for what neural networks are doing internally. One hypothesised cause of polysemanticity is superposition, where neural networks represent more features than they have neurons by assigning features to an overcomplete set of directions in activation space, rather than to individual neurons. Here, we attempt to identify those directions, using sparse autoencoders to reconstruct the internal activations of a language model. These autoencoders learn sets of sparsely activating features that are more interpretable and monosemantic than directions identified by alternative approaches, where interpretability is measured by automated methods. Ablating these features enables precise model editing, for example, by removing capabilities such as pronoun prediction, while disrupting model behaviour less than prior techniques. This work indicates that it is possible to resolve superposition in language models using a scalable, unsupervised method. Our method may serve as a foundation for future mechanistic interpretability work, which we hope will enable greater model transparency and steerability.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Sparse Autoencoders Find Highly Interpretable Features in Language Models | TensorX