TensorX
返回文献探索

Paper · arXiv 2503.14378

Impossible Videos

Zechen Bai, Hai Ci, Mike Zheng Shou

61 upvotesMarch 18, 2025arXiv 预印本
AI 摘要

IPV-Bench evaluates video generation and understanding models on creating and interpreting impossible videos, highlighting their limitations and guiding future advancements.

video generation modelsprompt followingcreativityvideo understanding modelstemporal dynamicsworld knowledgevideo-LLMsIPV-Bench

Abstract

Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts underexplored. This work aims to answer two questions: 1) Can today's video generation models effectively follow prompts to create impossible video content? 2) Are today's video understanding models good enough for understanding impossible videos? To this end, we introduce IPV-Bench, a novel benchmark designed to evaluate and foster progress in video understanding and generation. IPV-Bench is underpinned by a comprehensive taxonomy, encompassing 4 domains, 14 categories. It features diverse scenes that defy physical, biological, geographical, or social laws. Based on the taxonomy, a prompt suite is constructed to evaluate video generation models, challenging their prompt following and creativity capabilities. In addition, a video benchmark is curated to assess Video-LLMs on their ability of understanding impossible videos, which particularly requires reasoning on temporal dynamics and world knowledge. Comprehensive evaluations reveal limitations and insights for future directions of video models, paving the way for next-generation video models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号