TensorX
返回文献探索

Paper · arXiv 2608.29464

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini

15 upvotesAugust 29, 2026arXiv 预印本
AI 摘要

FACE-Eval reveals that chain-of-thought monitoring is less reliable when preference cues arrive via tool outputs or implicit artifacts, with lower verbalized commitment and higher hidden adoption across diverse open-weight models.

chain-of-thought monitoringfaithfulnessFACE-Evaltool-return cuesimplicit cuesverbalized commitmentunverbalized adoptionsource-attribution prompttranscript monitors

Abstract

Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples. We evaluate 15 open-weight models from eight families, with total parameters ranging from 4B to 1.60T. Every model has lower verbalized commitment for tool-return than user-message cues and for implicit than explicit cues. Unverbalized adoption is higher for tool-return cues on all 15 models and for implicit cues in 28 of 30 model-channel comparisons. A source-attribution prompt narrows the channel gap on seven models, sometimes by increasing user-channel unverbalized adoption, while telling models that their reasoning will be monitored does not reliably close the gap. We also use two transcript monitors (GPT-5.6-Luna and GPT-4o-mini) to detect preference adoption in the largest model of each family. Across 32 model-channel-explicitness cells, higher unverbalized adoption is associated with lower detection ability for both monitors (Pearson r=-0.54 and r=-0.78, respectively). These results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered | TensorX