TensorX
返回文献探索

Paper · arXiv 2305.16960

Training Socially Aligned Language Models in Simulated Human Society

Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M. Dai, Diyi Yang, Soroush Vosoughi

4 upvotesMay 26, 2023arXiv 预印本
AI 摘要

A novel training paradigm enables language models to learn from simulated social interactions, improving their alignment with societal norms and values compared to existing methods.

social alignmentAI systemslanguage modelstraining corpusadversarial attackssimulated social interactionsalignment benchmarks

Abstract

Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly replicate their training corpus in isolation, leading to subpar generalization in unfamiliar scenarios and vulnerability to adversarial attacks. This work presents a novel training paradigm that permits LMs to learn from simulated social interactions. In comparison to existing methodologies, our approach is considerably more scalable and efficient, demonstrating superior performance in alignment benchmarks and human evaluations. This paradigm shift in the training of LMs brings us a step closer to developing AI systems that can robustly and accurately reflect societal norms and values.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Training Socially Aligned Language Models in Simulated Human Society | TensorX