TensorX
返回文献探索

Paper · arXiv 2311.07961

The ART of LLM Refinement: Ask, Refine, and Trust

Kumar Shridhar, Koustuv Sinha, Andrew Cohen, Tianlu Wang, Ping Yu, Ram Pasunuru, Mrinmaya Sachan, Jason Weston, Asli Celikyilmaz

11 upvotesNovember 14, 2023arXiv 预印本
AI 摘要

The ART method improves the refinement of Large Language Model (LLM) outputs on multistep reasoning tasks by using smaller models to decide when and how to refine the initial predictions.

Large Language Modelsself-refinementreasoning with refinementARTmathematical word problemsquestion answeringGSM8KStrategyQAdecision makersmaller models

Abstract

In recent years, Large Language Models (LLMs) have demonstrated remarkable generative abilities, but can they judge the quality of their own generations? A popular concept, referred to as self-refinement, postulates that LLMs can detect and correct the errors in their generations when asked to do so. However, recent empirical evidence points in the opposite direction, suggesting that LLMs often struggle to accurately identify errors when reasoning is involved. To address this, we propose a reasoning with refinement objective called ART: Ask, Refine, and Trust, which asks necessary questions to decide when an LLM should refine its output, and either affirm or withhold trust in its refinement by ranking the refinement and the initial prediction. On two multistep reasoning tasks of mathematical word problems (GSM8K) and question answering (StrategyQA), ART achieves a performance gain of +5 points over self-refinement baselines, while using a much smaller model as the decision maker. We also demonstrate the benefit of using smaller models to make refinement decisions as a cost-effective alternative to fine-tuning a larger model.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
The ART of LLM Refinement: Ask, Refine, and Trust | TensorX