TensorX
返回文献探索

Paper · arXiv 2510.14972

TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar

Yinxi Li, Yuntian Deng, Pengyu Nie

35 upvotesOctober 16, 2025arXiv 预印本
AI 摘要

Misaligned tokenization in large language models for code leads to inconsistent model behavior, necessitating grammar-aware tokenization.

large language modelscode LLMssubword tokenizersbyte-pair encodingBPEsemantic-preserving rewrite rulesTokDrifttokenizationearly embeddingsgrammar token boundariescode understandingcode generation

Abstract

Large language models (LLMs) for code rely on subword tokenizers, such as byte-pair encoding (BPE), learned from mixed natural language text and programming language code but driven by statistics rather than grammar. As a result, semantically identical code snippets can be tokenized differently depending on superficial factors such as whitespace or identifier naming. To measure the impact of this misalignment, we introduce TokDrift, a framework that applies semantic-preserving rewrite rules to create code variants differing only in tokenization. Across nine code LLMs, including large ones with over 30B parameters, even minor formatting changes can cause substantial shifts in model behavior. Layer-wise analysis shows that the issue originates in early embeddings, where subword segmentation fails to capture grammar token boundaries. Our findings identify misaligned tokenization as a hidden obstacle to reliable code understanding and generation, highlighting the need for grammar-aware tokenization for future code LLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号