TensorX
返回文献探索

Paper · arXiv 2608.31022

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria

9 upvotesAugust 31, 2026arXiv 预印本
AI 摘要

The MNIST-PRO benchmark isolates perceptual-state construction in partially observable settings, revealing that multimodal agents struggle to integrate fragmented glimpses, continue exploring, and revise incorrect beliefs.

MNIST-PROactive sensingworking memoryperceptual stateglimpse-based searchpartial observabilitymultimodal modelsmemory representationsperceptual-state construction

Abstract

AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号