TensorX
返回文献探索

Paper · arXiv 2401.04081

MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts

Maciej Pióro, Kamil Ciebiera, Krystian Król, Jan Ludziejewski, Sebastian Jaszczur

74 upvotesJanuary 8, 2024arXiv 预印本
AI 摘要

A combination of Mixture of Experts (MoE) with State Space Models (SSMs) improves model efficiency and performance compared to both standalone SSMs and MoE-enhanced Transformers.

State Space ModelsSequential modelingTransformersMixture of ExpertsMoEMoE-Mamba

Abstract

State Space Models (SSMs) have become serious contenders in the field of sequential modeling, challenging the dominance of Transformers. At the same time, Mixture of Experts (MoE) has significantly improved Transformer-based LLMs, including recent state-of-the-art open-source models. We propose that to unlock the potential of SSMs for scaling, they should be combined with MoE. We showcase this on Mamba, a recent SSM-based model that achieves remarkable, Transformer-like performance. Our model, MoE-Mamba, outperforms both Mamba and Transformer-MoE. In particular, MoE-Mamba reaches the same performance as Mamba in 2.2x less training steps while preserving the inference performance gains of Mamba against the Transformer.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号