TensorX
返回文献探索

Paper · arXiv 2401.02330

LLaVA-φ: Efficient Multi-Modal Assistant with Small Language Model

Yichen Zhu, Minjie Zhu, Ning Liu, Zhicai Ou, Xiaofeng Mou, Jian Tang

18 upvotesJanuary 4, 2024arXiv 预印本
AI 摘要

A compact multi-modal assistant, LLaVA-Phi, using a smaller language model, Phi-2, achieves high performance in visual and multi-modal dialog tasks with resource efficiency.

multi-modal assistantPhi-2multi-modal dialoguesvisual comprehensionreasoningknowledge-based perceptionembodied agents

Abstract

In this paper, we introduce LLaVA-phi (LLaVA-Phi), an efficient multi-modal assistant that harnesses the power of the recently advanced small language model, Phi-2, to facilitate multi-modal dialogues. LLaVA-Phi marks a notable advancement in the realm of compact multi-modal models. It demonstrates that even smaller language models, with as few as 2.7B parameters, can effectively engage in intricate dialogues that integrate both textual and visual elements, provided they are trained with high-quality corpora. Our model delivers commendable performance on publicly available benchmarks that encompass visual comprehension, reasoning, and knowledge-based perception. Beyond its remarkable performance in multi-modal dialogue tasks, our model opens new avenues for applications in time-sensitive environments and systems that require real-time interaction, such as embodied agents. It highlights the potential of smaller language models to achieve sophisticated levels of understanding and interaction, while maintaining greater resource efficiency.The project is available at {https://github.com/zhuyiche/llava-phi}.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLaVA-φ: Efficient Multi-Modal Assistant with Small Language Model | TensorX