TensorX
返回文献探索

Paper · arXiv 2504.00050

JudgeLRM: Large Reasoning Models as a Judge

Nuo Chen, Zhiyuan Hu, Qingyun Zou, Jiaying Wu, Qian Wang, Bryan Hooi, Bingsheng He

61 upvotesMarch 31, 2025arXiv 预印本
AI 摘要

JudgeLRM models trained with reinforcement learning outperform existing models, especially in tasks requiring deep reasoning.

Large Language ModelsSupervised Fine-Tuningreinforcement learningjudge-wiseoutcome-driven rewardsF1 scoreDeepSeek-R1

Abstract

The rise of Large Language Models (LLMs) as evaluators offers a scalable alternative to human annotation, yet existing Supervised Fine-Tuning (SFT) for judges approaches often fall short in domains requiring complex reasoning. In this work, we investigate whether LLM judges truly benefit from enhanced reasoning capabilities. Through a detailed analysis of reasoning requirements across evaluation tasks, we reveal a negative correlation between SFT performance gains and the proportion of reasoning-demanding samples - highlighting the limitations of SFT in such scenarios. To address this, we introduce JudgeLRM, a family of judgment-oriented LLMs trained using reinforcement learning (RL) with judge-wise, outcome-driven rewards. JudgeLRM models consistently outperform both SFT-tuned and state-of-the-art reasoning models. Notably, JudgeLRM-3B surpasses GPT-4, and JudgeLRM-7B outperforms DeepSeek-R1 by 2.79% in F1 score, particularly excelling in judge tasks requiring deep reasoning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
JudgeLRM: Large Reasoning Models as a Judge | TensorX