TensorX
返回文献探索

Paper · arXiv 2601.00747

The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving

Max Ruiz Luyten, Mihaela van der Schaar

20 upvotesJanuary 2, 2026arXiv 预印本
AI 摘要

Large language model training methods that optimize for correctness can cause reasoning path diversity collapse, but a new variational framework provides principled solutions to maintain both accuracy and creativity.

large language modelsbootstrapped reasoning loopschains of thoughtreasoning pathssemantic entropyDistributional Creative Reasoningvariational objectivegradient flowprobability measuressolution tracesSTaRGRPODPOentropy bonuses

Abstract

State-of-the-art large language model (LLM) pipelines rely on bootstrapped reasoning loops: sampling diverse chains of thought and reinforcing the highest-scoring ones, mainly optimizing correctness. We analyze how this design choice is sensitive to the collapse of the model's distribution over reasoning paths, slashing semantic entropy and undermining creative problem-solving. To analyze this failure, we introduce Distributional Creative Reasoning (DCR), a unified variational objective that casts training as gradient flow through probability measures on solution traces. STaR, GRPO, and DPO, as well as entropy bonuses, and other methods, all constitute special cases of the same loss. The framework delivers three core results: (i) the diversity decay theorem, describing how correctness-based objectives lead to distinct modes of diversity decay for STaR, GRPO, and DPO; (ii) designs that ensure convergence to a stable and diverse policy, effectively preventing collapse; and (iii) simple, actionable recipes to achieve this in practice. DCR thus offers the first principled recipe for LLMs that remain both correct and creative.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号