TensorX
返回文献探索

Paper · arXiv 2604.11201

CocoaBench: Evaluating Unified Digital Agents in the Wild

CocoaBench Team, Shibo Hao, Zhining Zhang, Zhiqi Liang, Tianyang Liu, Yuheng Zha, Qiyue Gao, Jixuan Chen, Zilong Wang, Zhoujun Cheng, Haoxiang Zhang, Junli Wang, Hexi Jin, Boyuan Zheng, Kun Zhou, Yu Wang, Feng Yao, Licheng Liu, Yijiang Li, Zhifei Li, Zhengtao Han, Pracha Promthaw, Tommaso Cerruti, Xiaohan Fu, Ziqiao Ma, Jingbo Shang, Lianhui Qin, Julian McAuley, Eric P. Xing, Zhengzhong Liu, Rupesh Kumar Srivastava, Zhiting Hu

37 upvotesApril 13, 2026arXiv 预印本
AI 摘要

A new benchmark called CocoaBench evaluates unified digital agents on complex, multi-capability tasks requiring vision, search, and coding integration, revealing significant room for improvement in current agent systems.

LLM agentssoftware engineeringdeep researchGUI automationagent scaffoldsunified systemsdigital agentslong-horizon tasksvisionsearchcodingautomatic evaluationcontrolled comparisonmodel backbonesreasoningplanningtool useexecutionvisual grounding

Abstract

LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly integrating these capabilities into unified systems. Yet, most evaluations still test these capabilities in isolation, which leaves a gap for more diverse use cases that require agents to combine different capabilities. We introduce CocoaBench, a benchmark for unified digital agents built from human-designed, long-horizon tasks that require flexible composition of vision, search, and coding. Tasks are specified only by an instruction and an automatic evaluation function over the final output, enabling reliable and scalable evaluation across diverse agent infrastructures. We also present CocoaAgent, a lightweight shared scaffold for controlled comparison across model backbones. Experiments show that current agents remain far from reliable on CocoaBench, with the best evaluated system achieving only 45.1% success rate. Our analysis further points to substantial room for improvement in reasoning and planning, tool use and execution, and visual grounding.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
CocoaBench: Evaluating Unified Digital Agents in the Wild | TensorX