TensorX
返回文献探索

Paper · arXiv 2502.15007

LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev, Elizaveta Goncharova, Polina Druzhinina, Ivan Oseledets, Andrey Kuznetsov

175 upvotesFebruary 20, 2025arXiv 预印本
AI 摘要

Analysis of Large Language Models reveals that even minor tokens are crucial for context, and introduces LLM-Microscope for assessing token-level nonlinearity and contextual memory.

Large Language ModelsLLMscontextual informationtokensdeterminerspunctuationstopwordsarticlescommasMMLUBABILong-4kcontextualizationlinearityembeddingsLogit Lensintrinsic dimensionalitylong-range understanding

Abstract

We introduce methods to quantify how Large Language Models (LLMs) encode and store contextual information, revealing that tokens often seen as minor (e.g., determiners, punctuation) carry surprisingly high context. Notably, removing these tokens -- especially stopwords, articles, and commas -- consistently degrades performance on MMLU and BABILong-4k, even if removing only irrelevant tokens. Our analysis also shows a strong correlation between contextualization and linearity, where linearity measures how closely the transformation from one layer's embeddings to the next can be approximated by a single linear mapping. These findings underscore the hidden importance of filler tokens in maintaining context. For further exploration, we present LLM-Microscope, an open-source toolkit that assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions (via an adapted Logit Lens), and measures the intrinsic dimensionality of representations. This toolkit illuminates how seemingly trivial tokens can be critical for long-range understanding.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号