Safe RLHF: Safe Reinforcement Learning from Human Feedback
Josef Dai, Xuehai Pan, Ruiyang Sun +5 authors
Safe RLHF, a novel algorithm for human value alignment in LLMs, improves performance and safety by decoupling human preferences and using a Lagrangian method to balance reward and cost constraints.