TensorX
返回文献探索

Paper · arXiv 2501.04575

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection

Yuhang Liu, Pengxiang Li, Zishu Wei, Congkai Xie, Xueyu Hu, Xinchen Xu, Shengyu Zhang, Xiaotian Han, Hongxia Yang, Fei Wu

24 upvotesJanuary 8, 2025arXiv 预印本
AI 摘要

InfiGUIAgent, an MLLM-based GUI Agent, enhances multi-step reasoning and GUI interaction through a two-stage supervised fine-tuning pipeline.

multimodal large language modelsGUI Agentssupervised fine-tuningGUI understandinghierarchical reasoningexpectation-reflection reasoningsynthesized dataGUI benchmarks

Abstract

Graphical User Interface (GUI) Agents, powered by multimodal large language models (MLLMs), have shown great potential for task automation on computing devices such as computers and mobile phones. However, existing agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness. We introduce InfiGUIAgent, an MLLM-based GUI Agent trained with a two-stage supervised fine-tuning pipeline. Stage 1 enhances fundamental skills such as GUI understanding and grounding, while Stage 2 integrates hierarchical reasoning and expectation-reflection reasoning skills using synthesized data to enable native reasoning abilities of the agents. InfiGUIAgent achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks. Resources are available at https://github.com/Reallm-Labs/InfiGUIAgent.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号