TensorX
返回文献探索

Paper · arXiv 2307.03601

GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, Ping Luo

13 upvotesJuly 7, 2023arXiv 预印本
AI 摘要

GPT4RoI enhances vision-language multimodal understanding by enabling region-level alignment through instruction tuning, offering better control, capacity, and composition capabilities.

instruction tuningregion-of-interestspatial instructionvisual featureslanguage embeddingregion-level vision-language modelcontrollabilitysingle-regionmulti-regiondetailed region captioncomplex region reasoningobject detectorinformative object attributes

Abstract

Instruction tuning large language model (LLM) on image-text pairs has achieved unprecedented vision-language multimodal abilities. However, their vision-language alignments are only built on image-level, the lack of region-level alignment limits their advancements to fine-grained multimodal understanding. In this paper, we propose instruction tuning on region-of-interest. The key design is to reformulate the bounding box as the format of spatial instruction. The interleaved sequences of visual features extracted by the spatial instruction and the language embedding are input to LLM, and trained on the transformed region-text data in instruction tuning format. Our region-level vision-language model, termed as GPT4RoI, brings brand new conversational and interactive experience beyond image-level understanding. (1) Controllability: Users can interact with our model by both language and spatial instructions to flexibly adjust the detail level of the question. (2) Capacities: Our model supports not only single-region spatial instruction but also multi-region. This unlocks more region-level multimodal capacities such as detailed region caption and complex region reasoning. (3) Composition: Any off-the-shelf object detector can be a spatial instruction provider so as to mine informative object attributes from our model, like color, shape, material, action, relation to other objects, etc. The code, data, and demo can be found at https://github.com/jshilong/GPT4RoI.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest | TensorX