TensorX
返回文献探索

Paper · arXiv 2410.13785

PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment

Zekun Moore Wang, Shawn Wang, Kang Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, Wenhao Huang

19 upvotesOctober 17, 2024arXiv 预印本
AI 摘要

PopAlign enhances large language model alignment by introducing diverse contrasting patterns across prompts, models, and pipelines, leading to improved alignment compared to traditional methods.

large language modelsLLMsalignmentpreference-contrastive output pairsRLHFRLAIFcontrasting patternsjailbreaking attacksPopAligncomprehensive alignment

Abstract

Alignment of large language models (LLMs) involves training models on preference-contrastive output pairs to adjust their responses according to human preferences. To obtain such contrastive pairs, traditional methods like RLHF and RLAIF rely on limited contrasting patterns, such as varying model variants or decoding temperatures. This singularity leads to two issues: (1) alignment is not comprehensive; and thereby (2) models are susceptible to jailbreaking attacks. To address these issues, we investigate how to construct more comprehensive and diversified contrasting patterns to enhance preference data (RQ1) and verify the impact of the diversification of contrasting patterns on model alignment (RQ2). For RQ1, we propose PopAlign, a framework that integrates diversified contrasting patterns across the prompt, model, and pipeline levels, introducing six contrasting strategies that do not require additional feedback labeling procedures. Regarding RQ2, we conduct thorough experiments demonstrating that PopAlign significantly outperforms existing methods, leading to more comprehensive alignment.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment | TensorX