TensorX
返回文献探索

Paper · arXiv 2308.01399

Learning to Model the World with Language

Jessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner, Pieter Abbeel, Dan Klein, Anca Dragan

36 upvotesJuly 31, 2023arXiv 预印本
AI 摘要

Dynalang, a multimodal agent, learns to predict future text and image representations using language hints to improve task performance and enrich its understanding.

multimodal world modelself-supervised learningfuture predictionimagined model rolloutslanguage hintsgrid worldsphotorealistic scansenvironment descriptionsgame rulesinstructions

Abstract

To interact with humans in the world, agents need to understand the diverse types of language that people use, relate them to the visual world, and act based on them. While current agents learn to execute simple language instructions from task rewards, we aim to build agents that leverage diverse language that conveys general knowledge, describes the state of the world, provides interactive feedback, and more. Our key idea is that language helps agents predict the future: what will be observed, how the world will behave, and which situations will be rewarded. This perspective unifies language understanding with future prediction as a powerful self-supervised learning objective. We present Dynalang, an agent that learns a multimodal world model that predicts future text and image representations and learns to act from imagined model rollouts. Unlike traditional agents that use language only to predict actions, Dynalang acquires rich language understanding by using past language also to predict future language, video, and rewards. In addition to learning from online interaction in an environment, Dynalang can be pretrained on datasets of text, video, or both without actions or rewards. From using language hints in grid worlds to navigating photorealistic scans of homes, Dynalang utilizes diverse types of language to improve task performance, including environment descriptions, game rules, and instructions.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Learning to Model the World with Language | TensorX