TensorX
返回文献探索

Paper · arXiv 2504.05288

LiveVQA: Live Visual Knowledge Seeking

Mingyang Fu, Yuyang Peng, Benlin Liu, Yao Wan, Dongping Chen

15 upvotesApril 7, 2025arXiv 预印本
AI 摘要

Evaluation of various MLLMs on LiveVQA, a dataset of visual questions with up-to-date visual knowledge, shows that advanced visual reasoning is essential for complex multi-hop questions.

MLLMsGPT-4oGemma-3Qwen-2.5-VLvisual reasoningvisual questionsmulti-hop questionsvisual knowledge

Abstract

We introduce LiveVQA, an automatically collected dataset of latest visual knowledge from the Internet with synthesized VQA problems. LiveVQA consists of 3,602 single- and multi-hop visual questions from 6 news websites across 14 news categories, featuring high-quality image-text coherence and authentic information. Our evaluation across 15 MLLMs (e.g., GPT-4o, Gemma-3, and Qwen-2.5-VL family) demonstrates that stronger models perform better overall, with advanced visual reasoning capabilities proving crucial for complex multi-hop questions. Despite excellent performance on textual problems, models with tools like search engines still show significant gaps when addressing visual questions requiring latest visual knowledge, highlighting important areas for future research.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LiveVQA: Live Visual Knowledge Seeking | TensorX