TensorX
返回文献探索

Paper · arXiv 2601.11000

When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs

Zhongxiang Sun, Yi Zhan, Chenglei Shen, Weijie Yu, Xiao Zhang, Ming He, Jun Xu

27 upvotesJanuary 16, 2026arXiv 预印本
AI 摘要

Personalized large language models can generate false information aligned with user history instead of factual truth, but a new method called FPPS helps maintain both factual accuracy and personalized responses while preserving existing personalization effects.

personalized large language modelsfactual reasoningpersonalization-induced hallucinationsrepresentational entanglementFactuality-Preserving Personalized SteeringPFQABenchinference-time approachfactual accuracypersonalized performance

Abstract

Personalized large language models (LLMs) adapt model behavior to individual users to enhance user satisfaction, yet personalization can inadvertently distort factual reasoning. We show that when personalized LLMs face factual queries, there exists a phenomenon where the model generates answers aligned with a user's prior history rather than the objective truth, resulting in personalization-induced hallucinations that degrade factual reliability and may propagate incorrect beliefs, due to representational entanglement between personalization and factual representations. To address this issue, we propose Factuality-Preserving Personalized Steering (FPPS), a lightweight inference-time approach that mitigates personalization-induced factual distortions while preserving personalized behavior. We further introduce PFQABench, the first benchmark designed to jointly evaluate factual and personalized question answering under personalization. Experiments across multiple LLM backbones and personalization methods show that FPPS substantially improves factual accuracy while maintaining personalized performance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs | TensorX