TensorX
返回文献探索

Paper · arXiv 2408.02210

ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning

Yuxuan Wang, Alan Yuille, Zhuowan Li, Zilong Zheng

8 upvotesAugust 5, 2024arXiv 预印本
AI 摘要

An ExoViP method enhances vision-language programming by incorporating verification modules to correct planning and execution errors, improving compositional reasoning tasks.

compositional visual reasoninglarge language modelsfew-shot/zero-shot plannersvision-language programmingintrospective verificationverification modulesexoskeletonssub-verifiersvisual modulereasoning tracecompositional reasoning tasksstandard benchmarksopen-domain multi-modal challenges

Abstract

Compositional visual reasoning methods, which translate a complex query into a structured composition of feasible visual tasks, have exhibited a strong potential in complicated multi-modal tasks. Empowered by recent advances in large language models (LLMs), this multi-modal challenge has been brought to a new stage by treating LLMs as few-shot/zero-shot planners, i.e., vision-language (VL) programming. Such methods, despite their numerous merits, suffer from challenges due to LLM planning mistakes or inaccuracy of visual execution modules, lagging behind the non-compositional models. In this work, we devise a "plug-and-play" method, ExoViP, to correct errors in both the planning and execution stages through introspective verification. We employ verification modules as "exoskeletons" to enhance current VL programming schemes. Specifically, our proposed verification module utilizes a mixture of three sub-verifiers to validate predictions after each reasoning step, subsequently calibrating the visual module predictions and refining the reasoning trace planned by LLMs. Experimental results on two representative VL programming methods showcase consistent improvements on five compositional reasoning tasks on standard benchmarks. In light of this, we believe that ExoViP can foster better performance and generalization on open-domain multi-modal challenges.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning | TensorX