TensorX
返回文献探索

Paper · arXiv 2402.19427

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, Caglar Gulcehre

59 upvotesFebruary 29, 2024arXiv 预印本
AI 摘要

Hawk and Griffin, models incorporating gated linear recurrences, achieve high performance with lower computational resources and better hardware efficiency compared to existing architectures.

RNNsgated linear recurrenceslocal attentionMambaLlama-2tokensextrapolationhardware efficiencyTransformerssharding

Abstract

Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale. We propose Hawk, an RNN with gated linear recurrences, and Griffin, a hybrid model that mixes gated linear recurrences with local attention. Hawk exceeds the reported performance of Mamba on downstream tasks, while Griffin matches the performance of Llama-2 despite being trained on over 6 times fewer tokens. We also show that Griffin can extrapolate on sequences significantly longer than those seen during training. Our models match the hardware efficiency of Transformers during training, and during inference they have lower latency and significantly higher throughput. We scale Griffin up to 14B parameters, and explain how to shard our models for efficient distributed training.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models | TensorX