LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu, Hao Cheng, Haotian Liu +10 authors
LLaVA-Plus, a general-purpose multimodal assistant, enhances large multimodal models by integrating pre-trained vision and vision-language models, performing tool-assisted tasks and improving interaction through direct image grounding.