TensorX
返回文献探索

Paper · arXiv 2603.01562

RubricBench: Aligning Model-Generated Rubrics with Human Standards

Qiyuan Zhang, Junyi Zhou, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma

64 upvotesMarch 2, 2026arXiv 预印本
AI 摘要

RubricBench is introduced as a benchmark for evaluating rubric-guided reward models in large language model alignment, addressing the lack of discriminative complexity and ground-truth annotations in existing benchmarks.

Reward Modelsrubric-guided evaluationLarge Language Modelsbenchmarkpairwise comparisonsatomic rubricsmulti-dimensional filtration pipelinesurface-level biasesdiscriminative complexity

Abstract

As Large Language Model (LLM) alignment evolves from simple completions to complex, highly sophisticated generation, Reward Models are increasingly shifting toward rubric-guided evaluation to mitigate surface-level biases. However, the community lacks a unified benchmark to assess this evaluation paradigm, as existing benchmarks lack both the discriminative complexity and the ground-truth rubric annotations required for rigorous analysis. To bridge this gap, we introduce RubricBench, a curated benchmark with 1,147 pairwise comparisons specifically designed to assess the reliability of rubric-based evaluation. Our construction employs a multi-dimensional filtration pipeline to target hard samples featuring nuanced input complexity and misleading surface bias, augmenting each with expert-annotated, atomic rubrics derived strictly from instructions. Comprehensive experiments reveal a substantial capability gap between human-annotated and model-generated rubrics, indicating that even state-of-the-art models struggle to autonomously specify valid evaluation criteria, lagging considerably behind human-guided performance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
RubricBench: Aligning Model-Generated Rubrics with Human Standards | TensorX