TensorX
返回文献探索

Paper · arXiv 2306.13776

Swin-Free: Achieving Better Cross-Window Attention and Efficiency with Size-varying Window

Jinkyu Koo, John Yang, Le An, Gwenaelle Cunha Sergio, Su Inn Park

5 upvotesJune 23, 2023arXiv 预印本
AI 摘要

Swin-Free introduces size-varying windows to reduce memory copy operations, improving inference speed and accuracy over Swin Transformer in computer vision tasks.

Swin TransformerVision Transformerquadratic complexityself-attentionshifting windowsinferencecross-connectionlocal windowsSwin-Free

Abstract

Transformer models have shown great potential in computer vision, following their success in language tasks. Swin Transformer is one of them that outperforms convolution-based architectures in terms of accuracy, while improving efficiency when compared to Vision Transformer (ViT) and its variants, which have quadratic complexity with respect to the input size. Swin Transformer features shifting windows that allows cross-window connection while limiting self-attention computation to non-overlapping local windows. However, shifting windows introduces memory copy operations, which account for a significant portion of its runtime. To mitigate this issue, we propose Swin-Free in which we apply size-varying windows across stages, instead of shifting windows, to achieve cross-connection among local windows. With this simple design change, Swin-Free runs faster than the Swin Transformer at inference with better accuracy. Furthermore, we also propose a few of Swin-Free variants that are faster than their Swin Transformer counterparts.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Swin-Free: Achieving Better Cross-Window Attention and Efficiency with Size-varying Window | TensorX