TensorX
返回文献探索

Paper · arXiv 2404.01744

Octopus v2: On-device language model for super agent

Wei Chen, Zhiyuan Li

59 upvotesApril 2, 2024arXiv 预印本
AI 摘要

A new on-device method improves the performance of a large language model, matching GPT-4's accuracy and latency while reducing context length and enhancing deployment on edge devices.

function callinglarge-scale language modelson-device modelsGPT-4RAG-based function calling mechanismcontext lengthedge devices

Abstract

Language models have shown effectiveness in a variety of software applications, particularly in tasks related to automatic workflow. These models possess the crucial ability to call functions, which is essential in creating AI agents. Despite the high performance of large-scale language models in cloud environments, they are often associated with concerns over privacy and cost. Current on-device models for function calling face issues with latency and accuracy. Our research presents a new method that empowers an on-device model with 2 billion parameters to surpass the performance of GPT-4 in both accuracy and latency, and decrease the context length by 95\%. When compared to Llama-7B with a RAG-based function calling mechanism, our method enhances latency by 35-fold. This method reduces the latency to levels deemed suitable for deployment across a variety of edge devices in production environments, aligning with the performance requisites for real-world applications.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号