TensorX
返回文献探索

Paper · arXiv 2509.22638

Language Models Can Learn from Verbal Feedback Without Scalar Rewards

Renjie Luo, Zichen Liu, Xiangyan Liu, Chao Du, Min Lin, Wenhu Chen, Wei Lu, Tianyu Pang

70 upvotesSeptember 26, 2025arXiv 预印本
AI 摘要

Feedback-conditional policy (FCP) enables LLMs to learn from verbal feedback by treating it as a conditioning signal, improving expressiveness over scalar rewards.

LLMsRLhuman feedbackAI feedbackscalar rewardsfeedback-conditional policylanguage priorstext-to-image generationresponse-feedback pairsmaximum likelihood trainingoffline dataonline bootstrappingconditional generationreward optimization

Abstract

LLMs are often trained with RL from human or AI feedback, yet such methods typically compress nuanced feedback into scalar rewards, discarding much of their richness and inducing scale imbalance. We propose treating verbal feedback as a conditioning signal. Inspired by language priors in text-to-image generation, which enable novel outputs from unseen prompts, we introduce the feedback-conditional policy (FCP). FCP learns directly from response-feedback pairs, approximating the feedback-conditional posterior through maximum likelihood training on offline data. We further develop an online bootstrapping stage where the policy generates under positive conditions and receives fresh feedback to refine itself. This reframes feedback-driven learning as conditional generation rather than reward optimization, offering a more expressive way for LLMs to directly learn from verbal feedback. Our code is available at https://github.com/sail-sg/feedback-conditional-policy.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Language Models Can Learn from Verbal Feedback Without Scalar Rewards | TensorX