TensorX
返回文献探索

Paper · arXiv 2308.04592

Shepherd: A Critic for Language Model Generation

Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O'Brien, Ramakanth Pasunuru, Jane Dwivedi-Yu, Olga Golovneva, Luke Zettlemoyer, Maryam Fazel-Zarandi, Asli Celikyilmaz

33 upvotesAugust 8, 2023arXiv 预印本
AI 摘要

Shepherd, a small language model tuned for critique, outperforms or ties with larger models like ChatGPT in refining language model outputs using a high-quality feedback dataset.

language modelcritiquefeedback datasetGPT-4human evaluationChatGPTwin-rate

Abstract

As large language models improve, there is increasing interest in techniques that leverage these models' capabilities to refine their own outputs. In this work, we introduce Shepherd, a language model specifically tuned to critique responses and suggest refinements, extending beyond the capabilities of an untuned model to identify diverse errors and provide suggestions to remedy them. At the core of our approach is a high quality feedback dataset, which we curate from community feedback and human annotations. Even though Shepherd is small (7B parameters), its critiques are either equivalent or preferred to those from established models including ChatGPT. Using GPT-4 for evaluation, Shepherd reaches an average win-rate of 53-87% compared to competitive alternatives. In human evaluation, Shepherd strictly outperforms other models and on average closely ties with ChatGPT.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Shepherd: A Critic for Language Model Generation | TensorX