TensorX
返回文献探索

Paper · arXiv 2311.05437

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, Chunyuan Li

52 upvotesNovember 9, 2023arXiv 预印本
AI 摘要

LLaVA-Plus, a general-purpose multimodal assistant, enhances large multimodal models by integrating pre-trained vision and vision-language models, performing tool-assisted tasks and improving interaction through direct image grounding.

multimodal assistantpre-trained vision and vision-language modelsmultimodal instruction-following datatool usevisual understandinggenerationexternal knowledge retrievalcompositionshuman-AI interactiondirect image grounding

Abstract

LLaVA-Plus is a general-purpose multimodal assistant that expands the capabilities of large multimodal models. It maintains a skill repository of pre-trained vision and vision-language models and can activate relevant tools based on users' inputs to fulfill real-world tasks. LLaVA-Plus is trained on multimodal instruction-following data to acquire the ability to use tools, covering visual understanding, generation, external knowledge retrieval, and compositions. Empirical results show that LLaVA-Plus outperforms LLaVA in existing capabilities and exhibits new ones. It is distinct in that the image query is directly grounded and actively engaged throughout the entire human-AI interaction sessions, significantly improving tool use performance and enabling new scenarios.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents | TensorX