TensorX
返回文献探索

Paper · arXiv 2309.08637

TextBind: Multi-turn Interleaved Multimodal Instruction-following

Huayang Li, Siheng Li, Deng Cai, Longyue Wang, Lemao Liu, Taro Watanabe, Yujiu Yang, Shuming Shi

7 upvotesSeptember 14, 2023arXiv 预印本
AI 摘要

TextBind is an annotation-free framework that enhances large language models for multimodal instruction following by generating conversations from image-caption pairs.

large language modelsinstruction-following abilitiesgeneralizabilityreal-world tasksnatural language interfacesmultimodal instruction followingTextBindmulti-turn interleaved multimodal instruction-followingimage-caption pairsconversationsdatasetdemo

Abstract

Large language models with instruction-following abilities have revolutionized the field of artificial intelligence. These models show exceptional generalizability to tackle various real-world tasks through their natural language interfaces. However, their performance heavily relies on high-quality exemplar data, which is often difficult to obtain. This challenge is further exacerbated when it comes to multimodal instruction following. We introduce TextBind, an almost annotation-free framework for empowering larger language models with the multi-turn interleaved multimodal instruction-following capabilities. Our approach requires only image-caption pairs and generates multi-turn multimodal instruction-response conversations from a language model. We release our dataset, model, and demo to foster future research in the area of multimodal instruction following.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
TextBind: Multi-turn Interleaved Multimodal Instruction-following | TensorX