TensorX
返回文献探索

Paper · arXiv 2403.06098

VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

Wenhao Wang, Yi Yang

16 upvotesMarch 10, 2024arXiv 预印本
AI 摘要

A large-scale dataset of text-to-video prompts and generated videos opens new research avenues for improving text-to-video diffusion models.

TEXT-to-video diffusion modelsVidProMlarge-scale datasettext-to-video promptsvideo generationprompt engineeringvideo copy detection

Abstract

The arrival of Sora marks a new era for text-to-video diffusion models, bringing significant advancements in video generation and potential applications. However, Sora, as well as other text-to-video diffusion models, highly relies on the prompts, and there is no publicly available dataset featuring a study of text-to-video prompts. In this paper, we introduce VidProM, the first large-scale dataset comprising 1.67 million unique text-to-video prompts from real users. Additionally, the dataset includes 6.69 million videos generated by four state-of-the-art diffusion models and some related data. We initially demonstrate the curation of this large-scale dataset, which is a time-consuming and costly process. Subsequently, we show how the proposed VidProM differs from DiffusionDB, a large-scale prompt-gallery dataset for image generation. Based on the analysis of these prompts, we identify the necessity for a new prompt dataset specifically designed for text-to-video generation and gain insights into the preferences of real users when creating videos. Our large-scale and diverse dataset also inspires many exciting new research areas. For instance, to develop better, more efficient, and safer text-to-video diffusion models, we suggest exploring text-to-video prompt engineering, efficient video generation, and video copy detection for diffusion models. We make the collected dataset VidProM publicly available at GitHub and Hugging Face under the CC-BY- NC 4.0 License.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models | TensorX