TensorX
返回文献探索

Paper · arXiv 2410.21220

Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Zhixin Zhang, Yiyuan Zhang, Xiaohan Ding, Xiangyu Yue

11 upvotesOctober 28, 2024arXiv 预印本
AI 摘要

Vision Search Assistant enhances VLMs by collaborating with web agents for retrieving and generating information about unseen visual content, improving performance on image-based QA tasks.

vision-language modelsVLMsvisual understandingweb agentsopen-world Retrieval-Augmented Generationvisual and textual representationsopen-set QAclosed-set QA

Abstract

Search engines enable the retrieval of unknown information with texts. However, traditional methods fall short when it comes to understanding unfamiliar visual content, such as identifying an object that the model has never seen before. This challenge is particularly pronounced for large vision-language models (VLMs): if the model has not been exposed to the object depicted in an image, it struggles to generate reliable answers to the user's question regarding that image. Moreover, as new objects and events continuously emerge, frequently updating VLMs is impractical due to heavy computational burdens. To address this limitation, we propose Vision Search Assistant, a novel framework that facilitates collaboration between VLMs and web agents. This approach leverages VLMs' visual understanding capabilities and web agents' real-time information access to perform open-world Retrieval-Augmented Generation via the web. By integrating visual and textual representations through this collaboration, the model can provide informed responses even when the image is novel to the system. Extensive experiments conducted on both open-set and closed-set QA benchmarks demonstrate that the Vision Search Assistant significantly outperforms the other models and can be widely applied to existing VLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines | TensorX