TensorX
返回文献探索

Paper · arXiv 2412.06464

Gated Delta Networks: Improving Mamba2 with Delta Rule

Songlin Yang, Jan Kautz, Ali Hatamizadeh

17 upvotesDecember 9, 2024arXiv 预印本
AI 摘要

Gated DeltaNet, combining gating and delta update mechanisms, outperforms existing models in various tasks and achieves better performance through hybrid architectures with additional attention layers.

Linear Transformersstandard Transformersgatingadaptive memory controldelta update ruleparallel training algorithmGated DeltaNetMamba2DeltaNetlanguage modelingcommon-sense reasoningin-context retrievallength extrapolationlong-context understandingsliding window attention

Abstract

Linear Transformers have gained attention as efficient alternatives to standard Transformers, but their performance in retrieval and long-context tasks has been limited. To address these limitations, recent work has explored two distinct mechanisms: gating for adaptive memory control and the delta update rule for precise memory modifications. We observe that these mechanisms are complementary: gating enables rapid memory erasure while the delta rule facilitates targeted updates. Building on this insight, we introduce the gated delta rule and develop a parallel training algorithm optimized for modern hardware. Our proposed architecture, Gated DeltaNet, consistently surpasses existing models like Mamba2 and DeltaNet across multiple benchmarks, including language modeling, common-sense reasoning, in-context retrieval, length extrapolation, and long-context understanding. We further enhance performance by developing hybrid architectures that combine Gated DeltaNet layers with sliding window attention or Mamba2 layers, achieving both improved training efficiency and superior task performance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号