TensorX
返回文献探索

Paper · arXiv 2410.06961

Self-Boosting Large Language Models with Synthetic Preference Data

Qingxiu Dong, Li Dong, Xingxing Zhang, Zhifang Sui, Furu Wei

16 upvotesOctober 9, 2024arXiv 预印本
AI 摘要

SynPO, a self-boosting paradigm using synthetic preference data, enhances LLMs' instruction-following abilities and general performance without extensive human annotation.

Large Language ModelsLLMsSynPOself-boostingsynthetic preference dataself-prompt generatorresponse improvergenerative rewardsinstruction-following abilitiesAlpacaEvalArenaHardOpen LLM leaderboard

Abstract

Through alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Self-Boosting Large Language Models with Synthetic Preference Data | TensorX