TensorX
返回文献探索

Paper · arXiv 2502.03492

Teaching Language Models to Critique via Reinforcement Learning

Zhihui Xie, Jie chen, Liyu Chen, Weichao Mao, Jingjing Xu, Lingpeng Kong

24 upvotesFebruary 5, 2025arXiv 预印本
AI 摘要

A reinforcement learning framework trains a critic model to generate feedback for code generation, enhancing performance and mitigating errors.

LLMscode generationcritic modelreinforcement learningcorrection performancerelative improvementsgenerative reward modelsiterative critique-revision

Abstract

Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide accurate judgments and actionable suggestions. In this work, we study LLM critics for code generation and propose CTRL, a framework for Critic Training via Reinforcement Learning, which trains a critic model to generate feedback that maximizes correction performance for a fixed generator model without human supervision. Our results demonstrate that critics trained with CTRL significantly enhance pass rates and mitigate compounding errors across both base and stronger generator models. Furthermore, we show that these critic models act as accurate generative reward models and enable test-time scaling through iterative critique-revision, achieving up to 106.1% relative improvements across challenging code generation benchmarks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号