TensorX
返回文献探索

Paper · arXiv 2409.15268

Style over Substance: Failure Modes of LLM Judges in Alignment Benchmarking

Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, John P. Dickerson

12 upvotesSeptember 23, 2024arXiv 预印本
AI 摘要

SOS-Bench evaluates the effectiveness of preference optimization methods in aligning LLMs, finding that LLM-judgment preferences do not correlate well with safety, knowledge, and instruction following, and that supervised fine-tuning has the greatest impact on alignment.

preference optimizationLLM judgesSOS-BenchLLM meta-benchmarkworld knowledgeinstruction followingsupervised fine-tuningprompt diversity

Abstract

The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior alignment by virtue of better correspondence with human pairwise preferences, often measured by LLM judges. In this work, we attempt to answer the following question -- do LLM-judge preferences translate to progress on other, more concrete metrics for alignment, and if not, why not? We define a concrete metric for alignment, and introduce SOS-Bench, the largest standardized, reproducible LLM meta-benchmark to date. We find that (1) LLM-judgments do not correlate with concrete measures of safety, world knowledge, and instruction following; (2) LLM judges have powerful implicit biases, prioritizing style over factuality and safety; and (3) the supervised fine-tuning (SFT) stage of post-training, and not the PO stage, has the greatest impact on alignment, with data scaling and prompt diversity as the driving factors. Our codebase and complete results can be found at https://github.com/penfever/sos-bench.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Style over Substance: Failure Modes of LLM Judges in Alignment Benchmarking | TensorX