TensorX
返回文献探索

Paper · arXiv 2310.03744

Improved Baselines with Visual Instruction Tuning

Haotian Liu, Chunyuan Li, Yuheng Li, Yong Jae Lee

39 upvotesOctober 5, 2023arXiv 预印本
AI 摘要

Modifications to LLaVA using CLIP-ViT-L-336px and academic VQA data improve multimodal performance, achieving state-of-the-art results with limited training data.

LLaVACLIP-ViT-L-336pxMLP projectionVQAbenchmarks

Abstract

Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal connector in LLaVA is surprisingly powerful and data-efficient. With simple modifications to LLaVA, namely, using CLIP-ViT-L-336px with an MLP projection and adding academic-task-oriented VQA data with simple response formatting prompts, we establish stronger baselines that achieve state-of-the-art across 11 benchmarks. Our final 13B checkpoint uses merely 1.2M publicly available data, and finishes full training in ~1 day on a single 8-A100 node. We hope this can make state-of-the-art LMM research more accessible. Code and model will be publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号