TensorX
返回文献探索

Paper · arXiv 2504.21850

COMPACT: COMPositional Atomic-to-Complex Visual Capability Tuning

Xindi Wu, Hee Seung Hwang, Polina Kirichenko, Olga Russakovsky

27 upvotesApril 30, 2025arXiv 预印本
AI 摘要

COMPACT, a compositional visual capability tuning method, improves multimodal large language models' performance on complex vision-language tasks with less data than traditional visual instruction tuning.

multimodal large language modelsMLLMssimple vision-language taskscomplex tasksobject recognitioncountingspatial relationshipsVisual Instruction TuningVITcompositional complexityatomic capabilitiescomplex capabilitiesbenchmarksLLaVA-665kMMStarMM-Vet

Abstract

Multimodal Large Language Models (MLLMs) excel at simple vision-language tasks but struggle when faced with complex tasks that require multiple capabilities, such as simultaneously recognizing objects, counting them, and understanding their spatial relationships. This might be partially the result of the fact that Visual Instruction Tuning (VIT), a critical training step for MLLMs, has traditionally focused on scaling data volume, but not the compositional complexity of training examples. We propose COMPACT (COMPositional Atomic-to-complex visual Capability Tuning), which generates a training dataset explicitly controlling for the compositional complexity of the training examples. The data from COMPACT allows MLLMs to train on combinations of atomic capabilities to learn complex capabilities more efficiently. Across all benchmarks, COMPACT achieves comparable performance to the LLaVA-665k VIT while using less than 10% of its data budget, and even outperforms it on several, especially those involving complex multi-capability tasks. For example, COMPACT achieves substantial 83.3% improvement on MMStar and 94.0% improvement on MM-Vet compared to the full-scale VIT on particularly complex questions that require four or more atomic capabilities. COMPACT offers a scalable, data-efficient, visual compositional tuning recipe to improve on complex visual-language tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号