TensorX
返回文献探索

Paper · arXiv 2609.07816

Kalman Delta Networks: Uncertainty-aware Associative Memory

Ngoc Bui, Tinglin Huang, Rex Ying

27 upvotesSeptember 7, 2026arXiv 预印本
AI 摘要

Kalman Delta Networks reformulate linear attention as a linear-Gaussian state-space model with Kalman-filter updates to track memory uncertainty, yielding efficient scan-compatible approximations that improve language modeling performance.

linear attentionrecurrent associative memorylinear-Gaussian state-space modelKalman filterKalman Delta NetworksKalman gainDelta-ruleRiccati recursionDiagonal KDNIsotropic KDNassociative scansMobius maps

Abstract

Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号