TensorX
返回文献探索

Paper · arXiv 2307.15771

The Hydra Effect: Emergent Self-repair in Language Model Computations

Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, Shane Legg

20 upvotesJuly 28, 2023arXiv 预印本
AI 摘要

Analysis reveals adaptive computation and downregulation mechanisms in language models, showing loose coupling among layers even without dropout.

causal analysisadaptive computationattention layerHydra effectMLP layersmaximum-likelihood tokendownregulationcircuit-level attribution

Abstract

We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one attention layer of a language model cause another layer to compensate (which we term the Hydra effect) and (2) a counterbalancing function of late MLP layers that act to downregulate the maximum-likelihood token. Our ablation studies demonstrate that language model layers are typically relatively loosely coupled (ablations to one layer only affect a small number of downstream layers). Surprisingly, these effects occur even in language models trained without any form of dropout. We analyse these effects in the context of factual recall and consider their implications for circuit-level attribution in language models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号