TensorX
返回文献探索

Paper · arXiv 2506.06395

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Pengyi Li, Matvey Skripkin, Alexander Zubrey, Andrey Kuznetsov, Ivan Oseledets

135 upvotesJune 5, 2025arXiv 预印本
AI 摘要

Reinforcement Learning via Self-Confidence (RLSC) improves large language model accuracy using the model's confidence as a reward signal, eliminating the need for human labels or reward engineering.

Reinforcement LearningLarge language modelsself-confidenceRLSC

Abstract

Large language models (LLMs) excel at reasoning, yet post-training remains critical for aligning their behavior with task goals. Existing reinforcement learning (RL) methods often depend on costly human annotations or external reward models. We propose Reinforcement Learning via Self-Confidence (RLSC), which uses the model's own confidence as reward signals-eliminating the need for labels, preference models, or reward engineering. Applied to Qwen2.5-Math-7B with only 16 samples per question and 10 or 20 training steps, RLSC improves accuracy by +13.4% on AIME2024, +21.2% on MATH500, +21.7% on Minerva Math, +20.8% on Olympiadbench, and +9.7% on AMC23. RLSC provides a simple, scalable post-training method for inference models, requiring only a small number of samples and unlabelled supervision.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models | TensorX