TensorX
返回文献探索

Paper · arXiv 2507.15024

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback

Qiaoyu Tang, Hao Xiang, Le Yu, Bowen Yu, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Junyang Lin

14 upvotesJuly 20, 2025arXiv 预印本
AI 摘要

RefCritic, a reinforcement learning-based critic module with dual rule-based rewards, enhances model critique and refinement capabilities, demonstrating superior performance across multiple benchmarks compared to supervised fine-tuning methods.

Large Language Modelscritic modulessupervised fine-tuningreinforcement learningrule-based rewardsinstance-level correctnessrefinement accuraciesQwen2.5-14B-InstructDeepSeek-R1-Distill-Qwen-14BAIME25ProcessBenchmajority voting

Abstract

With the rapid advancement of Large Language Models (LLMs), developing effective critic modules for precise guidance has become crucial yet challenging. In this paper, we initially demonstrate that supervised fine-tuning for building critic modules (which is widely adopted in current solutions) fails to genuinely enhance models' critique abilities, producing superficial critiques with insufficient reflections and verifications. To unlock the unprecedented critique capabilities, we propose RefCritic, a long-chain-of-thought critic module based on reinforcement learning with dual rule-based rewards: (1) instance-level correctness of solution judgments and (2) refinement accuracies of the policy model based on critiques, aiming to generate high-quality evaluations with actionable feedback that effectively guides model refinement. We evaluate RefCritic on Qwen2.5-14B-Instruct and DeepSeek-R1-Distill-Qwen-14B across five benchmarks. On critique and refinement settings, RefCritic demonstrates consistent advantages across all benchmarks, e.g., 6.8\% and 7.2\% gains on AIME25 for the respective base models. Notably, under majority voting, policy models filtered by RefCritic show superior scaling with increased voting numbers. Moreover, despite training on solution-level supervision, RefCritic outperforms step-level supervised approaches on ProcessBench, a benchmark to identify erroneous steps in mathematical reasoning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号