SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Yao Zhao, Rishabh Joshi, Tianqi Liu +3 authors
Sequence Likelihood Calibration (SLiC) is shown to be an effective and simpler alternative to Reinforcement Learning from Human Feedback (RLHF) for learning from human preferences in language models.