Secrets of RLHF in Large Language Models Part I: PPO
Rui Zheng, Shihan Dou, Songyang Gao +24 authors
This report examines Reinfocement Learning with Human Feedback (RLHF) and proposes PPO-max to improve the stability of policy model training compared to other SFT models and ChatGPT.