TensorX
返回文献探索

Paper · arXiv 2603.17218

Alignment Makes Language Models Normative, Not Descriptive

Eilam Shapira, Moshe Tennenholtz, Roi Reichart

47 upvotesMarch 17, 2026arXiv 预印本
AI 摘要

Post-training alignment of language models creates a trade-off between human-like behavior prediction and normative performance, with base models better predicting complex strategic interactions while aligned models excel in simple, rule-based scenarios.

language modelshuman preference signalshuman behaviorstrategic gamesbargainingpersuasionnegotiationmatrix gamesnormative predictionsdescriptive dynamicsreciprocityretaliationhistory-dependent adaptation

Abstract

Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed human behavior. We compare 120 base-aligned model pairs on more than 10,000 real human decisions in multi-round strategic games - bargaining, persuasion, negotiation, and repeated matrix games. In these settings, base models outperform their aligned counterparts in predicting human choices by nearly 10:1, robustly across model families, prompt formulations, and game configurations. This pattern reverses, however, in settings where human behavior is more likely to follow normative predictions: aligned models dominate on one-shot textbook games across all 12 types tested and on non-strategic lottery choices - and even within the multi-round games themselves, at round one, before interaction history develops. This boundary-condition pattern suggests that alignment induces a normative bias: it improves prediction when human behavior is relatively well captured by normative solutions, but hurts prediction in multi-round strategic settings, where behavior is shaped by descriptive dynamics such as reciprocity, retaliation, and history-dependent adaptation. These results reveal a fundamental trade-off between optimizing models for human use and using them as proxies for human behavior.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Alignment Makes Language Models Normative, Not Descriptive | TensorX