TensorX
返回文献探索

Paper · arXiv 2603.28489

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

Muyang He, Hanzhong Guo, Junxiong Lin, Yizhou Yu

31 upvotesMarch 30, 2026arXiv 预印本
AI 摘要

Video generation models capable of simulating complex physical dynamics and long-horizon causality require efficient frameworks to become practical world simulators for interactive applications.

video generationworld simulatorsspatiotemporal modelingefficient modeling paradigmsefficient network architecturesefficient inference algorithmsautonomous drivingembodied AIgame simulation

Abstract

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between the theoretical capacity for world simulation and the heavy computational costs of spatiotemporal modeling. To address this, we comprehensively and systematically review video generation frameworks and techniques that consider efficiency as a crucial requirement for practical world modeling. We introduce a novel taxonomy in three dimensions: efficient modeling paradigms, efficient network architectures, and efficient inference algorithms. We further show that bridging this efficiency gap directly empowers interactive applications such as autonomous driving, embodied AI, and game simulation. Finally, we identify emerging research frontiers in efficient video-based world modeling, arguing that efficiency is a fundamental prerequisite for evolving video generators into general-purpose, real-time, and robust world simulators.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号