TensorX
返回文献探索

Paper · arXiv 2604.09574

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization

Jiachen Zhu, Lingyu Yang, Rong Shan, Congmin Zheng, Zeyu Zheng, Weiwen Liu, Yong Yu, Weinan Zhang, Jianghao Lin

31 upvotesFebruary 24, 2026arXiv 预印本
AI 摘要

Researchers propose humanization capabilities for autonomous GUI agents to avoid detection by digital platforms, introducing a benchmark and methods to balance imitability with task performance.

autonomous GUI agentsadversarial countermeasureshuman-centric ecosystemsTuring Test on ScreenMinMax optimizationLMM-based agentsbehavioral divergenceAgent Humanization Benchmarkimitabilitytask performance

Abstract

The rise of autonomous GUI agents has triggered adversarial countermeasures from digital platforms, yet existing research prioritizes utility and robustness over the critical dimension of anti-detection. We argue that for agents to survive in human-centric ecosystems, they must evolve Humanization capabilities. We introduce the ``Turing Test on Screen,'' formally modeling the interaction as a MinMax optimization problem between a detector and an agent aiming to minimize behavioral divergence. We then collect a new high-fidelity dataset of mobile touch dynamics, and conduct our analysis that vanilla LMM-based agents are easily detectable due to unnatural kinematics. Consequently, we establish the Agent Humanization Benchmark (AHB) and detection metrics to quantify the trade-off between imitability and utility. Finally, we propose methods ranging from heuristic noise to data-driven behavioral matching, demonstrating that agents can achieve high imitability theoretically and empirically without sacrificing performance. This work shifts the paradigm from whether an agent can perform a task to how it performs it within a human-centric ecosystem, laying the groundwork for seamless coexistence in adversarial digital environments.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号