TensorX
返回文献探索

Paper · arXiv 2306.04050

LLMZip: Lossless Text Compression using Large Language Models

Chandra Shekhara Kaushik Valmeekam, Krishna Narayanan, Dileep Kalathil, Jean-Francois Chamberland, Srinivas Shakkottai

5 upvotesJune 6, 2023arXiv 预印本
AI 摘要

A new estimate of English entropy using LLaMA-7B for next-token prediction leads to a compression algorithm that outperforms current text compression methods.

large language modelLLaMA-7Bnext-token predictionentropy estimationlossless compressionBSCZPAQpaq8h

Abstract

We provide new estimates of an asymptotic upper bound on the entropy of English using the large language model LLaMA-7B as a predictor for the next token given a window of past tokens. This estimate is significantly smaller than currently available estimates in cover1978convergent, lutati2023focus. A natural byproduct is an algorithm for lossless compression of English text which combines the prediction from the large language model with a lossless compression scheme. Preliminary results from limited experiments suggest that our scheme outperforms state-of-the-art text compression schemes such as BSC, ZPAQ, and paq8h.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLMZip: Lossless Text Compression using Large Language Models | TensorX