TensorX
返回文献探索

Paper · arXiv 2411.08790

Can sparse autoencoders be used to decompose and interpret steering vectors?

Harry Mayne, Yushi Yang, Adam Mahdi

8 upvotesNovember 13, 2024arXiv 预印本
AI 摘要

Investigation into why sparse autoencoders yield misleading decompositions of steering vectors in large language models reveals two main limitations.

sparse autoencoderssteering vectorsinput distributionnegative projectionsfeature directions

Abstract

Steering vectors are a promising approach to control the behaviour of large language models. However, their underlying mechanisms remain poorly understood. While sparse autoencoders (SAEs) may offer a potential method to interpret steering vectors, recent findings show that SAE-reconstructed vectors often lack the steering properties of the original vectors. This paper investigates why directly applying SAEs to steering vectors yields misleading decompositions, identifying two reasons: (1) steering vectors fall outside the input distribution for which SAEs are designed, and (2) steering vectors can have meaningful negative projections in feature directions, which SAEs are not designed to accommodate. These limitations hinder the direct use of SAEs for interpreting steering vectors.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Can sparse autoencoders be used to decompose and interpret steering vectors? | TensorX