TensorX
返回文献探索

Paper · arXiv 2402.17139

Video as the New Language for Real-World Decision Making

Sherry Yang, Jacob Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, Dale Schuurmans

21 upvotesFebruary 27, 2024arXiv 预印本
AI 摘要

Video generation can serve as a unified interface for diverse tasks and advance real-world AI applications through techniques like in-context learning and reinforcement learning, with opportunities in robotics, self-driving, and science.

in-context learningplanningreinforcement learning

Abstract

Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction. However, they have not been equally leveraged: language models have had significant real-world impact, whereas video generation has remained largely limited to media entertainment. Yet video data captures important information about the physical world that is difficult to express in language. To address this gap, we discuss an under-appreciated opportunity to extend video generation to solve tasks in the real world. We observe how, akin to language, video can serve as a unified interface that can absorb internet knowledge and represent diverse tasks. Moreover, we demonstrate how, like language models, video generation can serve as planners, agents, compute engines, and environment simulators through techniques such as in-context learning, planning and reinforcement learning. We identify major impact opportunities in domains such as robotics, self-driving, and science, supported by recent work that demonstrates how such advanced capabilities in video generation are plausibly within reach. Lastly, we identify key challenges in video generation that mitigate progress. Addressing these challenges will enable video generation models to demonstrate unique value alongside language models in a wider array of AI applications.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Video as the New Language for Real-World Decision Making | TensorX