TensorX
返回文献探索

Paper · arXiv 2412.14171

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, Saining Xie

24 upvotesDecember 18, 2024arXiv 预印本
AI 摘要

MLLMs trained on large video datasets show competitive but subhuman visual-spatial intelligence, with spatial reasoning as a key bottleneck that can be improved by generating cognitive maps.

Multimodal Large Language ModelsVSI-Benchvisual-spatial intelligencespatial reasoningcognitive maps

Abstract

Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language Models (MLLMs) trained on million-scale video datasets also ``think in space'' from videos? We present a novel video-based visual-spatial intelligence benchmark (VSI-Bench) of over 5,000 question-answer pairs, and find that MLLMs exhibit competitive - though subhuman - visual-spatial intelligence. We probe models to express how they think in space both linguistically and visually and find that while spatial reasoning capabilities remain the primary bottleneck for MLLMs to reach higher benchmark performance, local world models and spatial awareness do emerge within these models. Notably, prevailing linguistic reasoning techniques (e.g., chain-of-thought, self-consistency, tree-of-thoughts) fail to improve performance, whereas explicitly generating cognitive maps during question-answering enhances MLLMs' spatial distance ability.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces | TensorX