TensorX
返回文献探索

Paper · arXiv 2401.06102

Patchscope: A Unifying Framework for Inspecting Hidden Representations of Language Models

Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, Mor Geva

20 upvotesJanuary 11, 2024arXiv 预印本
AI 摘要

Patchscopes is a framework that uses LLMs to explain their internal representations, unifying and improving upon prior interpretability methods and enabling new applications like self-correction.

large language modelshidden representationsmodel behaviorPatchscopesinterpretability methodsvocabulary spaceearly layersexpressivityself-correctionmulti-hop reasoning

Abstract

Inspecting the information encoded in hidden representations of large language models (LLMs) can explain models' behavior and verify their alignment with human values. Given the capabilities of LLMs in generating human-understandable text, we propose leveraging the model itself to explain its internal representations in natural language. We introduce a framework called Patchscopes and show how it can be used to answer a wide range of research questions about an LLM's computation. We show that prior interpretability methods based on projecting representations into the vocabulary space and intervening on the LLM computation, can be viewed as special instances of this framework. Moreover, several of their shortcomings such as failure in inspecting early layers or lack of expressivity can be mitigated by a Patchscope. Beyond unifying prior inspection techniques, Patchscopes also opens up new possibilities such as using a more capable model to explain the representations of a smaller model, and unlocks new applications such as self-correction in multi-hop reasoning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Patchscope: A Unifying Framework for Inspecting Hidden Representations of Language Models | TensorX