TensorX
返回文献探索

Paper · arXiv 2310.20689

Learning From Mistakes Makes LLM Better Reasoner

Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, Jian-Guang Lou, Weizhu Chen

29 upvotesOctober 31, 2023arXiv 预印本
AI 摘要

LeMa, a learning-from-mistakes approach, enhances LLMs' mathematical reasoning by learning from inaccurate reasoning paths corrected by GPT-4, surpassing SOTA performance on math problems.

Large language modelsLearning from MistakesLeMaerror-driven learningfine-tuningmistake-correction data pairsinaccurate reasoning pathsGPT-4correctionfinal answerbackbone LLMsmathematical reasoning tasksWizardMathMetaMathGSM8KMATHpass@1 accuracySOTA

Abstract

Large language models (LLMs) recently exhibited remarkable reasoning capabilities on solving math problems. To further improve this capability, this work proposes Learning from Mistakes (LeMa), akin to human learning processes. Consider a human student who failed to solve a math problem, he will learn from what mistake he has made and how to correct it. Mimicking this error-driven learning process, LeMa fine-tunes LLMs on mistake-correction data pairs generated by GPT-4. Specifically, we first collect inaccurate reasoning paths from various LLMs and then employ GPT-4 as a "corrector" to (1) identify the mistake step, (2) explain the reason for the mistake, and (3) correct the mistake and generate the final answer. Experimental results demonstrate the effectiveness of LeMa: across five backbone LLMs and two mathematical reasoning tasks, LeMa consistently improves the performance compared with fine-tuning on CoT data alone. Impressively, LeMa can also benefit specialized LLMs such as WizardMath and MetaMath, achieving 85.4% pass@1 accuracy on GSM8K and 27.1% on MATH. This surpasses the SOTA performance achieved by non-execution open-source models on these challenging tasks. Our code, data and models will be publicly available at https://github.com/microsoft/CodeT.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号