TensorX
返回文献探索

Paper · arXiv 2407.02371

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, Ying Tai

56 upvotesJuly 2, 2024arXiv 预印本
AI 摘要

The paper introduces OpenVid-1M and OpenVidHD-0.4M datasets, along with MVDiT, a Multi-modal Video Diffusion Transformer, to enhance text-to-video generation by leveraging high-quality data and comprehensive text information.

Text-to-video generationSoraWebVid-10MPanda-70Mopen-scenario datasettext-video pairscross attention modulesemantic informationMulti-modal Video Diffusion TransformerMVDiTvisual tokenstext tokensvideo diffusionhigh-definition video generation

Abstract

Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previous popular video datasets, e.g. WebVid-10M and Panda-70M, are either with low quality or too large for most research institutions. Therefore, it is challenging but crucial to collect a precise high-quality text-video pairs for T2V generation. 2) Ignoring to fully utilize textual information. Recent T2V methods have focused on vision transformers, using a simple cross attention module for video generation, which falls short of thoroughly extracting semantic information from text prompt. To address these issues, we introduce OpenVid-1M, a precise high-quality dataset with expressive captions. This open-scenario dataset contains over 1 million text-video pairs, facilitating research on T2V generation. Furthermore, we curate 433K 1080p videos from OpenVid-1M to create OpenVidHD-0.4M, advancing high-definition video generation. Additionally, we propose a novel Multi-modal Video Diffusion Transformer (MVDiT) capable of mining both structure information from visual tokens and semantic information from text tokens. Extensive experiments and ablation studies verify the superiority of OpenVid-1M over previous datasets and the effectiveness of our MVDiT.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation | TensorX