TensorX
返回文献探索

Paper · arXiv 2505.10320

J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning

Chenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li, Jason Weston, Ilia Kulikov, Swarnadeep Saha

24 upvotesMay 15, 2025arXiv 预印本
AI 摘要

A reinforcement learning approach called J1 improves the judgment ability of LLM-as-a-Judge models through verifiable rewards and chain-of-thought reasoning.

reinforcement learningchain-of-thought reasoningjudgment tasksreward strategiesoffline trainingonline trainingPairwise-J1Pointwise-J1evaluation criteriaself-generated reference answersre-evaluation

Abstract

The progress of AI is bottlenecked by the quality of evaluation, and powerful LLM-as-a-Judge models have proved to be a core solution. Improved judgment ability is enabled by stronger chain-of-thought reasoning, motivating the need to find the best recipes for training such models to think. In this work we introduce J1, a reinforcement learning approach to training such models. Our method converts both verifiable and non-verifiable prompts to judgment tasks with verifiable rewards that incentivize thinking and mitigate judgment bias. In particular, our approach outperforms all other existing 8B or 70B models when trained at those sizes, including models distilled from DeepSeek-R1. J1 also outperforms o1-mini, and even R1 on some benchmarks, despite training a smaller model. We provide analysis and ablations comparing Pairwise-J1 vs Pointwise-J1 models, offline vs online training recipes, reward strategies, seed prompts, and variations in thought length and content. We find that our models make better judgments by learning to outline evaluation criteria, comparing against self-generated reference answers, and re-evaluating the correctness of model responses.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning | TensorX