ConvNets Match Vision Transformers at Scale
Samuel L. Smith, Andrew Brock, Leonard Berrada +1 authors
ConvNets pre-trained on a large dataset match the performance of Vision Transformers on ImageNet with comparable computational resources.
Trends · 研究趋势
数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。
Samuel L. Smith, Andrew Brock, Leonard Berrada +1 authors
ConvNets pre-trained on a large dataset match the performance of Vision Transformers on ImageNet with comparable computational resources.
Wenliang Dai, Junnan Li, Dongxu Li +6 authors
InstructBLIP models, built through instruction-aware visual feature extraction on BLIP-2, achieve top performance in vision-language tasks, including zero-shot evaluation and fine-tuned downstream tasks.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号