TensorX
返回文献探索

Paper · arXiv 2311.04498

NExT-Chat: An LMM for Chat, Detection and Segmentation

Ao Zhang, Liming Zhao, Chen-Wei Xie, Yun Zheng, Wei Ji, Tat-Seng Chua

13 upvotesNovember 8, 2023arXiv 预印本
AI 摘要

The pixel2emb method enables large multimodal models to effectively handle object location tasks by outputting location embeddings, leading to superior performance in multimodal conversations and various visual processing tasks.

large language models (LLMs)large multimodal models (LMMs)region-level understandingbounding box coordinatespixel2seqpixel2emblocation embeddingsmultimodal conversationsobject location modelingdetectionsegmentationvisual groundingregion captiongrounded reasoning

Abstract

The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance the level of visual comprehension, recent studies have equipped LMMs with region-level understanding capabilities by representing object bounding box coordinates as a series of text sequences (pixel2seq). In this paper, we introduce a novel paradigm for object location modeling called pixel2emb method, where we ask the LMM to output the location embeddings and then decoded by different decoders. This paradigm allows for different location formats (such as bounding boxes and masks) to be used in multimodal conversations Furthermore, this kind of embedding based location modeling enables the utilization of existing practices in localization tasks, such as detection and segmentation. In scenarios with limited resources, our pixel2emb demonstrates superior performance compared to existing state-of-the-art (SOTA) approaches in both the location input and output tasks under fair comparison. Leveraging the proposed pixel2emb method, we train an LMM named NExT-Chat and demonstrate its capability of handling multiple tasks like visual grounding, region caption, and grounded reasoning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号