TensorX
返回文献探索

Paper · arXiv 2408.02629

VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Zhiyu Tan, Xiaomeng Yang, Luozheng Qin, Hao Li

14 upvotesAugust 5, 2024arXiv 预印本
AI 摘要

A new, high-quality dataset called VidGen-1M improves training for text-to-video models by ensuring better video quality, detailed captions, and temporal consistency.

video-text pairstext-to-video modelstemporal consistencyvideo qualitycaption qualitydata imbalanceimage modelsmanual rule-based curationcoarse-to-fine curation strategyvideo generation model

Abstract

The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant shortcomings, including low temporal consistency, poor-quality captions, substandard video quality, and imbalanced data distribution. The prevailing video curation process, which depends on image models for tagging and manual rule-based curation, leads to a high computational load and leaves behind unclean data. As a result, there is a lack of appropriate training datasets for text-to-video models. To address this problem, we present VidGen-1M, a superior training dataset for text-to-video models. Produced through a coarse-to-fine curation strategy, this dataset guarantees high-quality videos and detailed captions with excellent temporal consistency. When used to train the video generation model, this dataset has led to experimental results that surpass those obtained with other models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation | TensorX