TensorX
返回文献探索

Paper · arXiv 2603.02765

Next Embedding Prediction Makes World Models Stronger

George Bredis, Nikita Balagansky, Daniil Gavrilov, Ruslan Rakhimov

21 upvotesMarch 3, 2026arXiv 预印本
AI 摘要

NE-Dreamer uses a temporal transformer to predict next-step encoder embeddings for model-based reinforcement learning without requiring decoders or auxiliary supervision.

temporal transformernext-step encoder embeddingstemporal predictive alignmentrepresentation spacemodel-based reinforcement learningDeepMind Control SuiteDMLabcoherent state representationspredictive state representations

Abstract

Capturing temporal dependencies is critical for model-based reinforcement learning (MBRL) in partially observable, high-dimensional domains. We introduce NE-Dreamer, a decoder-free MBRL agent that leverages a temporal transformer to predict next-step encoder embeddings from latent state sequences, directly optimizing temporal predictive alignment in representation space. This approach enables NE-Dreamer to learn coherent, predictive state representations without reconstruction losses or auxiliary supervision. On the DeepMind Control Suite, NE-Dreamer matches or exceeds the performance of DreamerV3 and leading decoder-free agents. On a challenging subset of DMLab tasks involving memory and spatial reasoning, NE-Dreamer achieves substantial gains. These results establish next-embedding prediction with temporal transformers as an effective, scalable framework for MBRL in complex, partially observable environments.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号