Training Language Models to Self-Correct via Reinforcement Learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal +15 authors
The multi-turn online reinforcement learning approach SCoRe improves LLM self-correction using self-generated data and regularization, achieving state-of-the-art performance.