TensorX
返回文献探索

Paper · arXiv 2407.06581

Vision language models are blind

Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, Anh Totti Nguyen

84 upvotesJuly 9, 2024arXiv 预印本
AI 摘要

State-of-the-art large language models with vision capabilities perform poorly on simple visual tasks, indicating limitations in their visual understanding.

large language modelsvision capabilitiesVLMsGPT-4oGemini 1.5 Provisual tasksvisual understanding benchmarksmyopiavisual detailsblindnesseducated guesses

Abstract

Large language models with vision capabilities (VLMs), e.g., GPT-4o and Gemini 1.5 Pro are powering countless image-text applications and scoring high on many vision-understanding benchmarks. Yet, we find that VLMs fail on 7 visual tasks absurdly easy to humans such as identifying (a) whether two circles overlap; (b) whether two lines intersect; (c) which letter is being circled in a word; and (d) counting the number of circles in a Olympic-like logo. The shockingly poor performance of four state-of-the-art VLMs suggests their vision is, at best, like of a person with myopia seeing fine details as blurry, and at worst, like an intelligent person that is blind making educated guesses. Code is available at: https://vlmsareblind.github.io/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号