TensorX
返回文献探索

Paper · arXiv 2507.09104

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

Taolin Zhang, Maosong Cao, Alexander Lam, Songyang Zhang, Kai Chen

18 upvotesJuly 12, 2025arXiv 预印本
AI 摘要

CompassJudger-2, a generalist judge model, improves cross-domain evaluation accuracy and robustness through task-driven data curation and a refined learning objective.

LLM-as-judgegeneralist judge modeltask-drivenmulti-domain data curationverifiable rewardsrejection samplingmargin policy gradient lossJudgerBenchV2cross-domain judgment accuracyrank consistency

Abstract

Recently, the role of LLM-as-judge in evaluating large language models has gained prominence. However, current judge models suffer from narrow specialization and limited robustness, undermining their capacity for comprehensive evaluations. In this work, we present CompassJudger-2, a novel generalist judge model that overcomes these limitations via a task-driven, multi-domain data curation strategy. Central to our approach is supervising judgment tasks with verifiable rewards, guiding intrinsic critical reasoning through rejection sampling to foster robust, generalizable judgment capabilities. We introduce a refined learning objective with margin policy gradient loss to enhance performance. Empirically, CompassJudger-2 achieves superior results across multiple judge and reward benchmarks, and our 7B model demonstrates competitive judgment accuracy with significantly larger models like DeepSeek-V3 and Qwen3-235B-A22B. Additionally, we propose JudgerBenchV2, a comprehensive benchmark evaluating cross-domain judgment accuracy and rank consistency to standardize judge model evaluation. These contributions advance robust, scalable LLM judgment and establish new performance and evaluation standards.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号