TensorX
返回文献探索

Paper · arXiv 2501.10132

ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario

Lucen Zhong, Zhengxiao Du, Xiaohan Zhang, Haiyi Hu, Jie Tang

22 upvotesJanuary 17, 2025arXiv 预印本
AI 摘要

ComplexFuncBench is a benchmark for evaluating complex function calling in large language models, including multi-step and constrained function calls, with an automatic evaluation framework called ComplexEval.

large language modelsfunction callingreal-world scenariosComplexFuncBenchComplexEvalmulti-stepconstrained function callinglong-parameter filingparameter value reasoninglong context

Abstract

Enhancing large language models (LLMs) with real-time APIs can help generate more accurate and up-to-date responses. However, evaluating the function calling abilities of LLMs in real-world scenarios remains under-explored due to the complexity of data collection and evaluation. In this work, we introduce ComplexFuncBench, a benchmark for complex function calling across five real-world scenarios. Compared to existing benchmarks, ComplexFuncBench encompasses multi-step and constrained function calling, which requires long-parameter filing, parameter value reasoning, and 128k long context. Additionally, we propose an automatic framework, ComplexEval, for quantitatively evaluating complex function calling tasks. Through comprehensive experiments, we demonstrate the deficiencies of state-of-the-art LLMs in function calling and suggest future directions for optimizing these capabilities. The data and code are available at https://github.com/THUDM/ComplexFuncBench.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario | TensorX