TensorX
返回文献探索

Paper · arXiv 2405.02246

What matters when building vision-language models?

Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor Sanh

104 upvotesMay 3, 2024arXiv 预印本
AI 摘要

Idefics2, a vision-language model with 8 billion parameters, achieves state-of-the-art performance on multimodal benchmarks through extensive experimental validation.

vision-language modelslarge language modelsvision transformerspre-trained modelsarchitecture choiceIdefics2multimodal benchmarks

Abstract

The growing interest in vision-language models (VLMs) has been driven by improvements in large language models and vision transformers. Despite the abundance of literature on this subject, we observe that critical decisions regarding the design of VLMs are often not justified. We argue that these unsupported decisions impede progress in the field by making it difficult to identify which choices improve model performance. To address this issue, we conduct extensive experiments around pre-trained models, architecture choice, data, and training methods. Our consolidation of findings includes the development of Idefics2, an efficient foundational VLM of 8 billion parameters. Idefics2 achieves state-of-the-art performance within its size category across various multimodal benchmarks, and is often on par with models four times its size. We release the model (base, instructed, and chat) along with the datasets created for its training.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号