TensorX
返回文献探索

Paper · arXiv 2306.16410

Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

William Berrios, Gautam Mittal, Tristan Thrush, Douwe Kiela, Amanpreet Singh

29 upvotesJune 28, 2023arXiv 预印本
AI 摘要

LENS uses language models to reason over outputs from vision modules, achieving competitive performance in vision and vision-language tasks without multimodal training.

LENSlarge language models (LLMs)vision moduleszero-shot object recognitionfew-shot object recognitionmultimodal training

Abstract

We propose LENS, a modular approach for tackling computer vision problems by leveraging the power of large language models (LLMs). Our system uses a language model to reason over outputs from a set of independent and highly descriptive vision modules that provide exhaustive information about an image. We evaluate the approach on pure computer vision settings such as zero- and few-shot object recognition, as well as on vision and language problems. LENS can be applied to any off-the-shelf LLM and we find that the LLMs with LENS perform highly competitively with much bigger and much more sophisticated systems, without any multimodal training whatsoever. We open-source our code at https://github.com/ContextualAI/lens and provide an interactive demo.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language | TensorX