TensorX
返回文献探索

Paper · arXiv 2406.09415

An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels

Duy-Kien Nguyen, Mahmoud Assran, Unnat Jain, Martin R. Oswald, Cees G. M. Snoek, Xinlei Chen

52 upvotesJune 13, 2024arXiv 预印本
AI 摘要

Vanilla Transformers achieve high performance in computer vision tasks by treating individual pixels as tokens, challenging the conventional use of local neighborhoods.

TransformerstokensVision TransformerConvNetslocal neighborhoodssupervised learningobject classificationself-supervised learningmasked autoencodingimage generationdiffusion models

Abstract

This work does not introduce a new method. Instead, we present an interesting finding that questions the necessity of the inductive bias -- locality in modern computer vision architectures. Concretely, we find that vanilla Transformers can operate by directly treating each individual pixel as a token and achieve highly performant results. This is substantially different from the popular design in Vision Transformer, which maintains the inductive bias from ConvNets towards local neighborhoods (e.g. by treating each 16x16 patch as a token). We mainly showcase the effectiveness of pixels-as-tokens across three well-studied tasks in computer vision: supervised learning for object classification, self-supervised learning via masked autoencoding, and image generation with diffusion models. Although directly operating on individual pixels is less computationally practical, we believe the community must be aware of this surprising piece of knowledge when devising the next generation of neural architectures for computer vision.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号