TensorX
返回文献探索

Paper · arXiv 2410.18558

Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Shuhao Gu, Jialing Zhang, Siyuan Zhou, Kevin Yu, Zhaohu Xing, Liangdong Wang, Zhou Cao, Jintao Jia, Zhuoyi Zhang, Yixuan Wang, Zhenchong Hu, Bo-Wen Zhang, Jijie Li, Dong Liang, Yingli Zhao, Yulong Ao, Yaoqi Liu, Fangxiang Feng, Guang Liu

19 upvotesOctober 24, 2024arXiv 预印本
AI 摘要

Infinity-MM, a large-scale multimodal instruction dataset, enhances open-source VLMs through rigorous filtering and synthetic data generation, achieving state-of-the-art performance in models of similar size.

Vision-Language Modelsmultimodal instruction datasetquality filteringdeduplicationsynthetic instruction generationAquila-VL-2Bstate-of-the-art performance

Abstract

Vision-Language Models (VLMs) have recently made significant progress, but the limited scale and quality of open-source instruction data hinder their performance compared to closed-source models. In this work, we address this limitation by introducing Infinity-MM, a large-scale multimodal instruction dataset with 40 million samples, enhanced through rigorous quality filtering and deduplication. We also propose a synthetic instruction generation method based on open-source VLMs, using detailed image annotations and diverse question generation. Using this data, we trained a 2-billion-parameter VLM, Aquila-VL-2B, achieving state-of-the-art (SOTA) performance for models of similar scale. This demonstrates that expanding instruction data and generating synthetic data can significantly improve the performance of open-source models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data | TensorX