TensorX
返回文献探索

Paper · arXiv 2503.04725

L^2M: Mutual Information Scaling Law for Long-Context Language Modeling

Zhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo, Marin Soljačić

21 upvotesMarch 6, 2025arXiv 预印本
AI 摘要

A scaling law for bipartite mutual information in natural language helps understand long-context language modeling and guides the development of large language models with longer context lengths.

bipartite mutual informationlong-range dependenciestwo-point mutual informationlong-context language modelingL^2M conditionlatent state sizetransformersstate space models

Abstract

We rigorously establish a bipartite mutual information scaling law in natural language that governs long-range dependencies. This scaling law, which we show is distinct from and scales independently of the conventional two-point mutual information, is the key to understanding long-context language modeling. Using this scaling law, we formulate the Long-context Language Modeling (L^2M) condition, which relates a model's capacity for effective long context length modeling to the scaling of its latent state size for storing past information. Our results are validated through experiments on both transformers and state space models. This work establishes a theoretical foundation that guides the development of large language models toward longer context lengths.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
L^2M: Mutual Information Scaling Law for Long-Context Language Modeling | TensorX