Self-Rewarding Language Models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho +3 authors
A study on Self-Rewarding Language Models shows that using LLM-as-a-Judge prompting for iterative DPO training enhances both instruction-following and self-reward generation, leading to superior performance compared to existing systems.