TensorX
返回文献探索

Paper · arXiv 2503.16416

Survey on Evaluation of LLM-based Agents

Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, Michal Shmueli-Scheuer

97 upvotesMarch 20, 2025arXiv 预印本
AI 摘要

This survey analyzes evaluation methodologies for large language model-based agents, covering fundamental capabilities, application-specific benchmarks, and generalist agents, highlighting trends and gaps in the field.

LLM-based agentsautonomous systemsevaluation methodologiesplanningtool useself-reflectionmemoryweb agentssoftware engineering agentsscientific agentsconversational agentsgeneralist agentsevaluation frameworksrealistic evaluationscontinuously updated benchmarkscost-efficiencysafetyrobustnessfine-grained evaluation methodsscalable evaluation methods

Abstract

The emergence of LLM-based agents represents a paradigm shift in AI, enabling autonomous systems to plan, reason, use tools, and maintain memory while interacting with dynamic environments. This paper provides the first comprehensive survey of evaluation methodologies for these increasingly capable agents. We systematically analyze evaluation benchmarks and frameworks across four critical dimensions: (1) fundamental agent capabilities, including planning, tool use, self-reflection, and memory; (2) application-specific benchmarks for web, software engineering, scientific, and conversational agents; (3) benchmarks for generalist agents; and (4) frameworks for evaluating agents. Our analysis reveals emerging trends, including a shift toward more realistic, challenging evaluations with continuously updated benchmarks. We also identify critical gaps that future research must address-particularly in assessing cost-efficiency, safety, and robustness, and in developing fine-grained, and scalable evaluation methods. This survey maps the rapidly evolving landscape of agent evaluation, reveals the emerging trends in the field, identifies current limitations, and proposes directions for future research.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Survey on Evaluation of LLM-based Agents | TensorX