TensorX
返回文献探索

Paper · arXiv 2307.12950

RLCD: Reinforcement Learning from Contrast Distillation for Language Model Alignment

Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, Yuandong Tian

11 upvotesJuly 24, 2023arXiv 预印本
AI 摘要

Reinforcement Learning from Contrast Distillation (RLCD) aligns language models to natural language principles using simulated preference pairs without human feedback, outperforming existing methods across various alignment tasks.

Reinforcement Learning from Contrast Distillation (RLCD)preference modelcontrasting positive and negative promptsreinforcement learningRLAIFcontext distillationalignment tasksharmlessnesshelpfulnessstory outline generation

Abstract

We propose Reinforcement Learning from Contrast Distillation (RLCD), a method for aligning language models to follow natural language principles without using human feedback. RLCD trains a preference model using simulated preference pairs that contain both a high-quality and low-quality example, generated using contrasting positive and negative prompts. The preference model is then used to improve a base unaligned language model via reinforcement learning. Empirically, RLCD outperforms RLAIF (Bai et al., 2022b) and context distillation (Huang et al., 2022) baselines across three diverse alignment tasks--harmlessness, helpfulness, and story outline generation--and on both 7B and 30B model scales for preference data simulation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
RLCD: Reinforcement Learning from Contrast Distillation for Language Model Alignment | TensorX