TensorX
返回文献探索

Paper · arXiv 2308.08316

Dual-Stream Diffusion Net for Text-to-Video Generation

Binhui Liu, Xin Liu, Anbo Dai, Zhiyong Zeng, Zhen Cui, Jian Yang

25 upvotesAugust 16, 2023arXiv 预印本
AI 摘要

The dual-stream diffusion net (DSDN) enhances video consistency and smoothness in text-to-video generation by using separate content and motion diffusion streams with a cross-transformer interaction module and motion decomposer/combiner.

dual-stream diffusion net (DSDN)diffusion streamsvideo contentmotion branchescross-transformer interaction modulemotion decomposermotion combiner

Abstract

With the emerging diffusion models, recently, text-to-video generation has aroused increasing attention. But an important bottleneck therein is that generative videos often tend to carry some flickers and artifacts. In this work, we propose a dual-stream diffusion net (DSDN) to improve the consistency of content variations in generating videos. In particular, the designed two diffusion streams, video content and motion branches, could not only run separately in their private spaces for producing personalized video variations as well as content, but also be well-aligned between the content and motion domains through leveraging our designed cross-transformer interaction module, which would benefit the smoothness of generated videos. Besides, we also introduce motion decomposer and combiner to faciliate the operation on video motion. Qualitative and quantitative experiments demonstrate that our method could produce amazing continuous videos with fewer flickers.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Dual-Stream Diffusion Net for Text-to-Video Generation | TensorX