TensorX
返回文献探索

Paper · arXiv 2310.16764

ConvNets Match Vision Transformers at Scale

Samuel L. Smith, Andrew Brock, Leonard Berrada, Soham De

21 upvotesOctober 25, 2023arXiv 预印本
AI 摘要

ConvNets pre-trained on a large dataset match the performance of Vision Transformers on ImageNet with comparable computational resources.

ConvNetsVision TransformersJFT-4Bpre-trainingTPU-v4NFNetlog-log scaling lawfine-tuningImageNetTop-1 accuracy

Abstract

Many researchers believe that ConvNets perform well on small or moderately sized datasets, but are not competitive with Vision Transformers when given access to datasets on the web-scale. We challenge this belief by evaluating a performant ConvNet architecture pre-trained on JFT-4B, a large labelled dataset of images often used for training foundation models. We consider pre-training compute budgets between 0.4k and 110k TPU-v4 core compute hours, and train a series of networks of increasing depth and width from the NFNet model family. We observe a log-log scaling law between held out loss and compute budget. After fine-tuning on ImageNet, NFNets match the reported performance of Vision Transformers with comparable compute budgets. Our strongest fine-tuned model achieves a Top-1 accuracy of 90.4%.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号