TensorX
返回文献探索

Paper · arXiv 2307.04087

SVIT: Scaling up Visual Instruction Tuning

Bo Zhao, Boya Wu, Tiejun Huang

7 upvotesJuly 9, 2023arXiv 预印本
AI 摘要

A new dataset of 3.2 million visual instruction tuning pairs improves multimodal performance in visual perception, reasoning, and planning by training on high-quality, diverse manual annotations.

foundation modelslarge language modelslarge vision modelsmultimodal abilityvisual captioningdialoguequestion answeringspeech instruction tuningSVITconversation question-answer pairscomplex reasoning question-answer pairsdetailed image descriptionsGPT-4visual perceptionvisual reasoningplanning

Abstract

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, dialogue, question answering, etc. Although existing multimodal models present impressive performance of visual understanding and reasoning, their limits are still largely under-explored due to the scarcity of high-quality instruction tuning data. To push the limits of multimodal capability, we Sale up Visual Instruction Tuning (SVIT) by constructing a dataset of 3.2 million visual instruction tuning data including 1.6M conversation question-answer (QA) pairs and 1.6M complex reasoning QA pairs and 106K detailed image descriptions. Besides the volume, the proposed dataset is also featured by the high quality and rich diversity, which is generated by prompting GPT-4 with the abundant manual annotations of images. We empirically verify that training multimodal models on SVIT can significantly improve the multimodal performance in terms of visual perception, reasoning and planing.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SVIT: Scaling up Visual Instruction Tuning | TensorX