TensorX
返回文献探索

Paper · arXiv 2501.10573

The Geometry of Tokens in Internal Representations of Large Language Models

Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Cazzaniga, Matteo Biagetti

9 upvotesJanuary 17, 2025arXiv 预印本
AI 摘要

The research explores the relationship between token embedding geometry and next token prediction in transformers, finding correlations between geometric properties and prediction loss.

transformer modelstoken embeddingsempirical measureintrinsic dimensionneighborhood overlapcosine similaritycross-entropy loss

Abstract

We investigate the relationship between the geometry of token embeddings and their role in the next token prediction within transformer models. An important aspect of this connection uses the notion of empirical measure, which encodes the distribution of token point clouds across transformer layers and drives the evolution of token representations in the mean-field interacting picture. We use metrics such as intrinsic dimension, neighborhood overlap, and cosine similarity to observationally probe these empirical measures across layers. To validate our approach, we compare these metrics to a dataset where the tokens are shuffled, which disrupts the syntactic and semantic structure. Our findings reveal a correlation between the geometric properties of token embeddings and the cross-entropy loss of next token predictions, implying that prompts with higher loss values have tokens represented in higher-dimensional spaces.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
The Geometry of Tokens in Internal Representations of Large Language Models | TensorX