TensorX
返回文献探索

Paper · arXiv 2507.13546

nablaNABLA: Neighborhood Adaptive Block-Level Attention

Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva, Vladimir Arkhipkin, Vladimir Korviakov, Vladimir Polovnikov, Viacheslav Vasilev, Evelina Sidorova, Denis Dimitrov

126 upvotesJuly 17, 2025arXiv 预印本
AI 摘要

NABLA, a Neighborhood Adaptive Block-Level Attention mechanism, enhances video diffusion transformers by reducing computational overhead without significantly impacting generative quality or visual fidelity.

transformer-based architecturesvideo generationfull attention mechanismsquadratic complexityNeighborhood Adaptive Block-Level Attentionvideo diffusion transformersblock-wise attentionadaptive sparsity-driven thresholdFlex Attention operatorCLIP scoreVBench scorehuman evaluation score

Abstract

Recent progress in transformer-based architectures has demonstrated remarkable success in video generation tasks. However, the quadratic complexity of full attention mechanisms remains a critical bottleneck, particularly for high-resolution and long-duration video sequences. In this paper, we propose NABLA, a novel Neighborhood Adaptive Block-Level Attention mechanism that dynamically adapts to sparsity patterns in video diffusion transformers (DiTs). By leveraging block-wise attention with adaptive sparsity-driven threshold, NABLA reduces computational overhead while preserving generative quality. Our method does not require custom low-level operator design and can be seamlessly integrated with PyTorch's Flex Attention operator. Experiments demonstrate that NABLA achieves up to 2.7x faster training and inference compared to baseline almost without compromising quantitative metrics (CLIP score, VBench score, human evaluation score) and visual quality drop. The code and model weights are available here: https://github.com/gen-ai-team/Wan2.1-NABLA

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号