TensorX
返回文献探索

Paper · arXiv 2403.13802

ZigMa: Zigzag Mamba Diffusion Model

Vincent Tao Hu, Stefan Andreas Baumann, Ming Gui, Olga Grebenkova, Pingchuan Ma, Johannes Fischer, Bjorn Ommer

18 upvotesMarch 20, 2024arXiv 预印本
AI 摘要

Zigzag Mamba, an improvement over Mamba, enhances visual data generation with better spatial continuity, speed, and memory usage, and is scalable with Stochastic Interpolant on large-resolution datasets.

diffusion modelscalabilityquadratic complexityState-Space ModelMambaspatial continuityscan schemeZigzag Mambazero-parameter methodStochastic Interpolantlarge-resolution visual datasetsFacesHQUCF101MultiModal-CelebA-HQMS COCO

Abstract

The diffusion model has long been plagued by scalability and quadratic complexity issues, especially within transformer-based structures. In this study, we aim to leverage the long sequence modeling capability of a State-Space Model called Mamba to extend its applicability to visual data generation. Firstly, we identify a critical oversight in most current Mamba-based vision methods, namely the lack of consideration for spatial continuity in the scan scheme of Mamba. Secondly, building upon this insight, we introduce a simple, plug-and-play, zero-parameter method named Zigzag Mamba, which outperforms Mamba-based baselines and demonstrates improved speed and memory utilization compared to transformer-based baselines. Lastly, we integrate Zigzag Mamba with the Stochastic Interpolant framework to investigate the scalability of the model on large-resolution visual datasets, such as FacesHQ 1024times 1024 and UCF101, MultiModal-CelebA-HQ, and MS COCO 256times 256. Code will be released at https://taohu.me/zigma/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ZigMa: Zigzag Mamba Diffusion Model | TensorX