TensorX
返回文献探索

Paper · arXiv 2311.10775

ToolTalk: Evaluating Tool-Usage in a Conversational Setting

Nicholas Farn, Richard Shin

9 upvotesNovember 15, 2023arXiv 预印本
AI 摘要

ToolTalk is a benchmark for evaluating LLM-based assistants using complex user intents requiring multi-step tool interactions, with success rates of 26% and 50% for GPT-3.5 and GPT-4 respectively.

Large language modelsLLMsToolTalkbenchmarkuser intentsmulti-step tool usagedialoguetoolspluginssimulated implementationexecution feedbackGPT-3.5GPT-4

Abstract

Large language models (LLMs) have displayed massive improvements in reason- ing and decision-making skills and can hold natural conversations with users. Many recent works seek to augment LLM-based assistants with external tools so they can access private or up-to-date information and carry out actions on behalf of users. To better measure the performance of these assistants, this paper introduces ToolTalk, a benchmark consisting of complex user intents re- quiring multi-step tool usage specified through dialogue. ToolTalk contains 28 tools grouped into 7 plugins, and includes a complete simulated implementa- tion of each tool, allowing for fully automated evaluation of assistants that rely on execution feedback. ToolTalk also emphasizes tools that externally affect the world rather than only tools for referencing or searching information. We evaluate GPT-3.5 and GPT-4 on ToolTalk resulting in success rates of 26% and 50% respectively. Our analysis of the errors reveals three major categories and suggests some future directions for improvement. We release ToolTalk at https://github.com/microsoft/ToolTalk.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号