TensorX
返回文献探索

Paper · arXiv 2601.15282

Rethinking Video Generation Model for the Embodied World

Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, Daquan Zhou

46 upvotesJanuary 21, 2026arXiv 预印本
AI 摘要

A comprehensive robotics benchmark evaluates video generation models across multiple task domains and robot embodiments, revealing significant gaps in physical realism and introducing a large-scale dataset to address training data limitations.

video generation modelsembodied intelligencerobotics benchmarkrobot-oriented video generationtask domainsphysical plausibilityaction completenessSpearman correlation coefficientRoVid-Xdata pipelinerobotic datasetvideo modelsembodied AIgeneral intelligence

Abstract

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing high-quality videos that accurately reflect real-world robotic interactions remains challenging, and the lack of a standardized benchmark limits fair comparisons and progress. To address this gap, we introduce a comprehensive robotics benchmark, RBench, designed to evaluate robot-oriented video generation across five task domains and four distinct embodiments. It assesses both task-level correctness and visual fidelity through reproducible sub-metrics, including structural consistency, physical plausibility, and action completeness. Evaluation of 25 representative models highlights significant deficiencies in generating physically realistic robot behaviors. Furthermore, the benchmark achieves a Spearman correlation coefficient of 0.96 with human evaluations, validating its effectiveness. While RBench provides the necessary lens to identify these deficiencies, achieving physical realism requires moving beyond evaluation to address the critical shortage of high-quality training data. Driven by these insights, we introduce a refined four-stage data pipeline, resulting in RoVid-X, the largest open-source robotic dataset for video generation with 4 million annotated video clips, covering thousands of tasks and enriched with comprehensive physical property annotations. Collectively, this synergistic ecosystem of evaluation and data establishes a robust foundation for rigorous assessment and scalable training of video models, accelerating the evolution of embodied AI toward general intelligence.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Rethinking Video Generation Model for the Embodied World | TensorX