TensorX
返回文献探索

Paper · arXiv 2409.09269

Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types

Neelabh Sinha, Vinija Jain, Aman Chadha

8 upvotesSeptember 14, 2024arXiv 预印本
AI 摘要

A framework for evaluating Vision-Language Models in VQA tasks using a novel dataset and GoEval metric, highlighting performance variation across models and guiding selection based on task and resource constraints.

Vision-Language ModelsVQAzero-shot inferenceevaluation frameworkmultimodal evaluation metricGoEvalGPT-4otask typesapplication domainsknowledge typesGemini-1.5-ProGPT-4o-miniInternVL-2-8BCogVLM-2-Llama-3-19B

Abstract

Visual Question-Answering (VQA) has become a key use-case in several applications to aid user experience, particularly after Vision-Language Models (VLMs) achieving good results in zero-shot inference. But evaluating different VLMs for an application requirement using a standardized framework in practical settings is still challenging. This paper introduces a comprehensive framework for evaluating VLMs tailored to VQA tasks in practical settings. We present a novel dataset derived from established VQA benchmarks, annotated with task types, application domains, and knowledge types, three key practical aspects on which tasks can vary. We also introduce GoEval, a multimodal evaluation metric developed using GPT-4o, achieving a correlation factor of 56.71% with human judgments. Our experiments with ten state-of-the-art VLMs reveals that no single model excelling universally, making appropriate selection a key design decision. Proprietary models such as Gemini-1.5-Pro and GPT-4o-mini generally outperform others, though open-source models like InternVL-2-8B and CogVLM-2-Llama-3-19B demonstrate competitive strengths in specific contexts, while providing additional advantages. This study guides the selection of VLMs based on specific task requirements and resource constraints, and can also be extended to other vision-language tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号