TensorX
返回文献探索

Paper · arXiv 2501.05727

Enabling Scalable Oversight via Self-Evolving Critic

Zhengyang Tang, Ziniu Li, Zhenyang Xiao, Tian Ding, Ruoyu Sun, Benyou Wang, Dayiheng Liu, Fei Huang, Tianyu Liu, Bowen Yu, Junyang Lin

72 upvotesJanuary 10, 2025arXiv 预印本
AI 摘要

SCRIT, a self-evolving critique framework, enhances LLMs' critique capabilities using synthetic data and self-validation, achieving significant improvements in critique-correction and error identification benchmarks.

Large Language Models (LLMs)self-evolvingsynthetic datacontrastive-based self-criticreference solutionsstep-by-step critiqueself-validationQwen2.5-72B-Instructcritique-correctionerror identification benchmarks

Abstract

Despite their remarkable performance, the development of Large Language Models (LLMs) faces a critical challenge in scalable oversight: providing effective feedback for tasks where human evaluation is difficult or where LLMs outperform humans. While there is growing interest in using LLMs for critique, current approaches still rely on human annotations or more powerful models, leaving the issue of enhancing critique capabilities without external supervision unresolved. We introduce SCRIT (Self-evolving CRITic), a framework that enables genuine self-evolution of critique abilities. Technically, SCRIT self-improves by training on synthetic data, generated by a contrastive-based self-critic that uses reference solutions for step-by-step critique, and a self-validation mechanism that ensures critique quality through correction outcomes. Implemented with Qwen2.5-72B-Instruct, one of the most powerful LLMs, SCRIT achieves up to a 10.3\% improvement on critique-correction and error identification benchmarks. Our analysis reveals that SCRIT's performance scales positively with data and model size, outperforms alternative approaches, and benefits critically from its self-validation component.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号