TensorX
返回文献探索

Paper · arXiv 2311.00571

LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Wei-Ge Chen, Irina Spiridonova, Jianwei Yang, Jianfeng Gao, Chunyuan Li

43 upvotesNovember 1, 2023arXiv 预印本
AI 摘要

LLaVA-Interactive is a cost-efficient multimodal dialogue system that integrates visual chat, image segmentation, and generation capabilities to enhance human-AI interaction.

multimodal human-AI interactiondialoguesmultimodal inputsvisual promptpre-built AI modelsvisual chatimage segmentationimage generationimage editingmultimodal skills

Abstract

LLaVA-Interactive is a research prototype for multimodal human-AI interaction. The system can have multi-turn dialogues with human users by taking multimodal user inputs and generating multimodal responses. Importantly, LLaVA-Interactive goes beyond language prompt, where visual prompt is enabled to align human intents in the interaction. The development of LLaVA-Interactive is extremely cost-efficient as the system combines three multimodal skills of pre-built AI models without additional model training: visual chat of LLaVA, image segmentation from SEEM, as well as image generation and editing from GLIGEN. A diverse set of application scenarios is presented to demonstrate the promises of LLaVA-Interactive and to inspire future research in multimodal interactive systems.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing | TensorX