TensorX
返回文献探索

Paper · arXiv 2507.19478

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents

Xuehui Wang, Zhenyu Wu, JingJing Xie, Zichen Ding, Bowen Yang, Zehao Li, Zhaoyang Liu, Qingyun Li, Xuan Dong, Zhe Chen, Weiyun Wang, Xiangyu Zhao, Jixuan Chen, Haodong Duan, Tianbao Xie, Chenyu Yang, Shiqian Su, Yue Yu, Yuan Huang, Yiqian Liu, Xiao Zhang, Yanting Zhang, Xiangyu Yue, Weijie Su, Xizhou Zhu, Wei Shen, Jifeng Dai, Wenhai Wang

33 upvotesJuly 25, 2025arXiv 预印本
AI 摘要

MMBench-GUI evaluates GUI automation agents across multiple platforms using a hierarchical benchmark and Efficiency-Quality Area metric, highlighting the importance of visual grounding, task planning, and efficiency.

GUI Content UnderstandingElement GroundingTask AutomationTask CollaborationEfficiency-Quality Area (EQA)visual groundingmodular frameworkstask planningcross-platform generalizationlong-context memorylong-term reasoningtask efficiencyprecise localizationeffective planningearly stopping strategies

Abstract

We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of modular frameworks that integrate specialized grounding modules. Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role. More important, task efficiency remains a critically underexplored dimension, and all models suffer from substantial inefficiencies, with excessive redundant steps even when tasks are ultimately completed. The integration of precise localization, effective planning, and early stopping strategies is indispensable to enable truly efficient and scalable GUI automation. Our benchmark code, evaluation data, and running environment will be publicly available at https://github.com/open-compass/MMBench-GUI.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents | TensorX