TensorX
返回文献探索

Paper · arXiv 2609.06289

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi

28 upvotesSeptember 5, 2026arXiv 预印本
AI 摘要

Activation steering vectors in large language models encode theory-aligned human value geometry when derived via distribution-driven methods, with geometric fidelity scaling with model size but declining after instruction tuning.

activation steeringsteering vectorslatent geometrySchwartz's Theory of Basic Human Valuesdistribution-driven methodsCAASphericalSteerODESteerCOLD-SteerBiPOinstruction tuning

Abstract

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman ρ up to 0.51, p < 10^{-13}). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号