TensorX
返回文献探索

Paper · arXiv 2405.21060

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Tri Dao, Albert Gu

69 upvotesMay 31, 2024arXiv 预印本
AI 摘要

A new framework connects state-space models and transformers, leading to a faster architecture for language modeling.

Transformersstate-space modelsMambaselective SSMstate space dualitystructured semiseparable matricesMamba-2

Abstract

While Transformers have been the main architecture behind deep learning's success in language modeling, state-space models (SSMs) such as Mamba have recently been shown to match or outperform Transformers at small to medium scale. We show that these families of models are actually quite closely related, and develop a rich framework of theoretical connections between SSMs and variants of attention, connected through various decompositions of a well-studied class of structured semiseparable matrices. Our state space duality (SSD) framework allows us to design a new architecture (Mamba-2) whose core layer is an a refinement of Mamba's selective SSM that is 2-8X faster, while continuing to be competitive with Transformers on language modeling.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号