TensorX
返回文献探索

Paper · arXiv 2306.04387

M^3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, Lingpeng Kong, Qi Liu

10 upvotesJune 7, 2023arXiv 预印本
AI 摘要

The M$^3$IT dataset and Ying-VLM model enhance vision-language models' alignment with human instructions across various languages and modalities.

M$^3$IT datasetvision-language modelsinstruction tuningvision-to-texttask coverageinstruction numberinstance scaleYing-VLMworld knowledgeunseen video tasks

Abstract

Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks. However, progress in open vision-language models (VLMs) has been limited due to the scarcity of high-quality instruction datasets. To tackle this challenge and promote research in the vision-language field, we introduce the Multi-Modal, Multilingual Instruction Tuning (M^3IT) dataset, designed to optimize VLM alignment with human instructions. Our M^3IT dataset comprises 40 carefully curated datasets, including 2.4 million instances and 400 manually written task instructions, reformatted into a vision-to-text structure. Key tasks are translated into 80 languages with an advanced translation system, ensuring broader accessibility. M^3IT surpasses previous datasets regarding task coverage, instruction number and instance scale. Moreover, we develop Ying-VLM, a VLM model trained on our M^3IT dataset, showcasing its potential to answer complex questions requiring world knowledge, generalize to unseen video tasks, and comprehend unseen instructions in Chinese. To encourage further research, we have open-sourced both the dataset and trained models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
M^3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning | TensorX