TensorX
返回文献探索

Paper · arXiv 2410.02884

LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Di Zhang, Jianbo Wu, Jingdi Lei, Tong Che, Jiatong Li, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone, Yuqiang Li, Wanli Ouyang, Dongzhan Zhou

53 upvotesOctober 3, 2024arXiv 预印本
AI 摘要

LLaMA-Berry enhances LLMs for mathematical reasoning by combining Monte Carlo Tree Search with Self-Refine and leveraging a Pairwise Preference Reward Model to optimize search efficiency and problem-solving capability.

Monte Carlo Tree SearchSelf-RefineSR-MCTSPairwise Preference Reward ModelRLHFEnhanced Borda CountGPQAAIME24AMC23ToTrStar

Abstract

This paper presents an advanced mathematical problem-solving framework, LLaMA-Berry, for enhancing the mathematical reasoning ability of Large Language Models (LLMs). The framework combines Monte Carlo Tree Search (MCTS) with iterative Self-Refine to optimize the reasoning path and utilizes a pairwise reward model to evaluate different paths globally. By leveraging the self-critic and rewriting capabilities of LLMs, Self-Refine applied to MCTS (SR-MCTS) overcomes the inefficiencies and limitations of conventional step-wise and greedy search algorithms by fostering a more efficient exploration of solution spaces. Pairwise Preference Reward Model~(PPRM), inspired by Reinforcement Learning from Human Feedback (RLHF), is then used to model pairwise preferences between solutions, utilizing an Enhanced Borda Count (EBC) method to synthesize these preferences into a global ranking score to find better answers. This approach addresses the challenges of scoring variability and non-independent distributions in mathematical reasoning tasks. The framework has been tested on general and advanced benchmarks, showing superior performance in terms of search efficiency and problem-solving capability compared to existing methods like ToT and rStar, particularly in complex Olympiad-level benchmarks, including GPQA, AIME24 and AMC23.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning | TensorX