TensorX
返回文献探索

Paper · arXiv 2501.09038

Do generative video models learn physical principles from watching videos?

Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, Robert Geirhos

33 upvotesJanuary 14, 2025arXiv 预印本
AI 摘要

A new benchmark dataset, Physics-IQ, shows that current AI video models achieve visual realism without understanding underlying physical principles, highlighting the need for further advancements.

Physics-IQSoraRunwayPikaLumiereStable Video DiffusionVideoPoetfluid dynamicsopticssolid mechanicsmagnetismthermodynamics

Abstract

AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn ``world models'' that discover laws of physics -- or, alternatively, are they merely sophisticated pixel predictors that achieve visual realism without understanding the physical principles of reality? We address this question by developing Physics-IQ, a comprehensive benchmark dataset that can only be solved by acquiring a deep understanding of various physical principles, like fluid dynamics, optics, solid mechanics, magnetism and thermodynamics. We find that across a range of current models (Sora, Runway, Pika, Lumiere, Stable Video Diffusion, and VideoPoet), physical understanding is severely limited, and unrelated to visual realism. At the same time, some test cases can already be successfully solved. This indicates that acquiring certain physical principles from observation alone may be possible, but significant challenges remain. While we expect rapid advances ahead, our work demonstrates that visual realism does not imply physical understanding. Our project page is at https://physics-iq.github.io; code at https://github.com/google-deepmind/physics-IQ-benchmark.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Do generative video models learn physical principles from watching videos? | TensorX