TensorX
返回文献探索

Paper · arXiv 2607.20734

LLMs Get Lost in Evolving User Intent

Jihoon Tack, Philippe Laban, Jennifer Neville

25 upvotesJuly 22, 2026arXiv 预印本
AI 摘要

Large language models show significant performance drops when user intent evolves across multi-turn conversations, revealing a critical gap in tracking dynamic goals.

LLMsmulti-turn conversationevolving intentcollaborative agentsstatic evaluationdynamic interaction

Abstract

As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLMs Get Lost in Evolving User Intent | TensorX