TensorX
返回文献探索

Paper · arXiv 2404.05902

WILBUR: Adaptive In-Context Learning for Robust and Accurate Web Agents

Michael Lutz, Arth Bohra, Manvel Saroyan, Artem Harutyunyan, Giovanni Campagna

22 upvotesApril 8, 2024arXiv 预印本
AI 摘要

Wilbur uses a differentiable ranking model and instruction synthesis to fine-tune a large language model for web agent tasks, achieving state-of-the-art results with text-only inputs and superior performance on the WebVoyager benchmark.

differentiable ranking modelinstruction synthesisgenerative auto-curriculumlarge language modelprompt populationintelligent backtrackingstate-of-the-artWebVoyager benchmarkmulti-modal model

Abstract

In the realm of web agent research, achieving both generalization and accuracy remains a challenging problem. Due to high variance in website structure, existing approaches often fail. Moreover, existing fine-tuning and in-context learning techniques fail to generalize across multiple websites. We introduce Wilbur, an approach that uses a differentiable ranking model and a novel instruction synthesis technique to optimally populate a black-box large language model's prompt with task demonstrations from previous runs. To maximize end-to-end success rates, we also propose an intelligent backtracking mechanism that learns and recovers from its mistakes. Finally, we show that our ranking model can be trained on data from a generative auto-curriculum which samples representative goals from an LLM, runs the agent, and automatically evaluates it, with no manual annotation. Wilbur achieves state-of-the-art results on the WebVoyager benchmark, beating text-only models by 8% overall, and up to 36% on certain websites. On the same benchmark, Wilbur is within 5% of a strong multi-modal model despite only receiving textual inputs, and further analysis reveals a substantial number of failures are due to engineering challenges of operating the web.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
WILBUR: Adaptive In-Context Learning for Robust and Accurate Web Agents | TensorX