TensorX
返回文献探索

Paper · arXiv 2603.08436

Can Vision-Language Models Solve the Shell Game?

Tiedong Liu, Wee Sun Lee

39 upvotesMarch 9, 2026arXiv 预印本
AI 摘要

Vision-Language Models exhibit poor performance on visual entity tracking due to reliance on static features; a proposed method using spatiotemporal grounded chain-of-thought achieves high accuracy by generating object trajectories as intermediate states.

Vision-Language Modelsspatiotemporal continuityvisual entity trackingtransformer-based VLMschain-of-thoughtobject trajectoriesintermediate supervisionexpressivity constraints

Abstract

Visual entity tracking is an innate cognitive ability in humans, yet it remains a critical bottleneck for Vision-Language Models (VLMs). This deficit is often obscured in existing video benchmarks by visual shortcuts. We introduce VET-Bench, a synthetic diagnostic testbed featuring visually identical objects that necessitate tracking exclusively through spatiotemporal continuity. Our experiments reveal that current state-of-the-art VLMs perform at or near chance level on VET-Bench, exposing a fundamental limitation: an over-reliance on static frame-level features and a failure to maintain entity representations over time. We provide a theoretical analysis drawing connections to the state-tracking problem, proving that fixed-depth transformer-based VLMs are fundamentally limited in tracking indistinguishable objects without intermediate supervision due to expressivity constraints. To address this, we propose Spatiotemporal Grounded Chain-of-Thought (SGCoT): generating object trajectories as explicit intermediate states. Leveraging Molmo2's object tracking ability, we elicit SGCoT reasoning by fine-tuning on synthesized text-only data for alignment. Our method achieves state-of-the-art accuracy exceeding 90% on VET-Bench, demonstrating that VLMs can reliably solve the video shell-game task end-to-end without external tools. Our code and data are available at https://vetbench.github.io .

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Can Vision-Language Models Solve the Shell Game? | TensorX