TensorX
返回文献探索

Paper · arXiv 2309.00071

YaRN: Efficient Context Window Extension of Large Language Models

Bowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico Shippole

86 upvotesAugust 31, 2023arXiv 预印本
AI 摘要

YaRN extends the context window of transformer-based language models like LLaMA with improved efficiency and performance.

Rotary Position EmbeddingsRoPEcompute-efficientcontext windowLLaMAfine-tuningcheckpoints

Abstract

Rotary Position Embeddings (RoPE) have been shown to effectively encode positional information in transformer-based language models. However, these models fail to generalize past the sequence length they were trained on. We present YaRN (Yet another RoPE extensioN method), a compute-efficient method to extend the context window of such models, requiring 10x less tokens and 2.5x less training steps than previous methods. Using YaRN, we show that LLaMA models can effectively utilize and extrapolate to context lengths much longer than their original pre-training would allow, while also surpassing previous the state-of-the-art at context window extension. In addition, we demonstrate that YaRN exhibits the capability to extrapolate beyond the limited context of a fine-tuning dataset. We publish the checkpoints of Llama 2 7B/13B fine-tuned using YaRN with 64k and 128k context windows at https://github.com/jquesnelle/yarn

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号