TensorX
返回文献探索

Paper · arXiv 2412.05939

Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models

Xiao Xu, Tianhao Niu, Yuxi Xie, Libo Qin, Wanxiang Che, Min-Yen Kan

15 upvotesDecember 8, 2024arXiv 预印本
AI 摘要

Integrating fine-grained concept annotations enhances multimodal large language models, improving their performance in vision-language tasks by aligning vision and language at multiple granularities.

Multimodal Large Language ModelsMLLMsfine-grained concept annotationsMultimodal Multi-Grained Concept annotationsMMGiCmultimodal comprehensionmultimodal generationPOPESEED-Bench

Abstract

Multimodal Large Language Models (MLLMs) excel in vision--language tasks by pre-training solely on coarse-grained concept annotations (e.g., image captions). We hypothesize that integrating fine-grained concept annotations (e.g., object labels and object regions) will further improve performance, as both data granularities complement each other in terms of breadth and depth in concept representation. We introduce a new dataset featuring Multimodal Multi-Grained Concept annotations (MMGiC) for MLLMs. In constructing MMGiC, we explore the impact of different data recipes on multimodal comprehension and generation. Our analyses reveal that multi-grained concept annotations integrate and complement each other, under our structured template and a general MLLM framework. We clearly explore and demonstrate the potential of MMGiC to help MLLMs better locate and learn concepts, aligning vision and language at multiple granularities. We further validate our hypothesis by investigating the fair comparison and effective collaboration between MMGiC and image--caption data on 12 multimodal comprehension and generation benchmarks, e.g., their appropriate combination achieve 3.95% and 2.34% absolute improvements over image--caption data alone on POPE and SEED-Bench. Code, data and models will be available at https://github.com/LooperXX/MMGiC.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号