TensorX
返回文献探索

Paper · arXiv 2503.15417

Temporal Regularization Makes Your Video Generator Stronger

Harold Haodong Chen, Haojian Huang, Xianfeng Wu, Yexin Liu, Yajing Bai, Wen-Jie Shu, Harry Yang, Ser-Nam Lim

22 upvotesMarch 19, 2025arXiv 预印本
AI 摘要

FluxFlow, a temporal augmentation strategy, enhances the temporal coherence and diversity in video generation without modifying model architectures.

temporal augmentationtemporal perturbationsspatial fidelityUCF-101VBenchU-NetDiTAR-based architectures

Abstract

Temporal quality is a critical aspect of video generation, as it ensures consistent motion and realistic dynamics across frames. However, achieving high temporal coherence and diversity remains challenging. In this work, we explore temporal augmentation in video generation for the first time, and introduce FluxFlow for initial investigation, a strategy designed to enhance temporal quality. Operating at the data level, FluxFlow applies controlled temporal perturbations without requiring architectural modifications. Extensive experiments on UCF-101 and VBench benchmarks demonstrate that FluxFlow significantly improves temporal coherence and diversity across various video generation models, including U-Net, DiT, and AR-based architectures, while preserving spatial fidelity. These findings highlight the potential of temporal augmentation as a simple yet effective approach to advancing video generation quality.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Temporal Regularization Makes Your Video Generator Stronger | TensorX