TensorX
返回文献探索

Paper · arXiv 2410.12381

HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks

Fengji Zhang, Linquan Wu, Huiyu Bai, Guancheng Lin, Xiao Li, Xiao Yu, Yue Wang, Bei Chen, Jacky Keung

43 upvotesOctober 16, 2024arXiv 预印本
AI 摘要

HumanEval-V provides a comprehensive benchmark for evaluating LMMs' diagram interpretation and reasoning abilities, particularly in coding contexts, revealing significant areas for improvement.

Large Multimodal ModelsHumanEval-Vdiagram interpretationvisual reasoningcoding contextsfunction signaturestest casescode generationspatial transformationstopological relationshipsdynamic patterns

Abstract

Understanding and reasoning over diagrams is a fundamental aspect of human intelligence. While Large Multimodal Models (LMMs) have demonstrated impressive capabilities across various tasks, existing benchmarks lack comprehensive evaluation of their diagram interpretation and reasoning abilities, particularly in coding contexts. We present HumanEval-V, a rigorous benchmark of human-annotated coding tasks that spans six task types and evaluates diverse visual reasoning capabilities. Each task features carefully crafted diagrams paired with function signatures and test cases, employing novel code generation tasks to thoroughly assess models' diagram comprehension. Through extensive experiments with 22 LMMs, we find that even top-performing models achieve modest success rates, with Claude 3.5 Sonnet reaching only 36.8% pass@1, highlighting substantial room for improvement. Our analysis reveals that current LMMs struggle with spatial transformations, topological relationships, and dynamic patterns that humans find intuitive. These findings provide valuable insights for advancing LMMs' visual reasoning abilities. We have open-sourced our code and benchmark at https://github.com/HumanEval-V/HumanEval-V-Benchmark.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号