TensorX
返回文献探索

Paper · arXiv 2308.00436

SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning

Ning Miao, Yee Whye Teh, Tom Rainforth

24 upvotesAugust 1, 2023arXiv 预印本
AI 摘要

A zero-shot verification scheme improves LLMs' question-answering performance by identifying and correcting errors in step-by-step reasoning.

large language models (LLMs)chain-of-thoughts (CoT)reasoning problemsnon-linear thinkingmulti-step reasoningzero-shot verification schemeweighted votingquestion-answeringGSM8KMathQAMATH

Abstract

The recent progress in large language models (LLMs), especially the invention of chain-of-thoughts (CoT) prompting, makes it possible to solve reasoning problems. However, even the strongest LLMs are still struggling with more complicated problems that require non-linear thinking and multi-step reasoning. In this work, we explore whether LLMs have the ability to recognize their own errors, without resorting to external resources. In particular, we investigate whether they can be used to identify individual errors within a step-by-step reasoning. To this end, we propose a zero-shot verification scheme to recognize such errors. We then use this verification scheme to improve question-answering performance, by using it to perform weighted voting on different generated answers. We test the method on three math datasets-GSM8K, MathQA, and MATH-and find that it successfully recognizes errors and, in turn, increases final predictive performance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号