TensorX
返回文献探索

Paper · arXiv 2501.08316

Diffusion Adversarial Post-Training for One-Step Video Generation

Shanchuan Lin, Xin Xia, Yuxi Ren, Ceyuan Yang, Xuefeng Xiao, Lu Jiang

36 upvotesJanuary 14, 2025arXiv 预印本
AI 摘要

Adversarial Post-Training improves the quality and speed of one-step video and image generation using diffusion models.

diffusion modelsone-step generationadversarial post-trainingR1 regularizationSeaweed-APTreal-time video generation

Abstract

The diffusion models are widely used for image and video generation, but their iterative generation process is slow and expansive. While existing distillation approaches have demonstrated the potential for one-step generation in the image domain, they still suffer from significant quality degradation. In this work, we propose Adversarial Post-Training (APT) against real data following diffusion pre-training for one-step video generation. To improve the training stability and quality, we introduce several improvements to the model architecture and training procedures, along with an approximated R1 regularization objective. Empirically, our experiments show that our adversarial post-trained model, Seaweed-APT, can generate 2-second, 1280x720, 24fps videos in real time using a single forward evaluation step. Additionally, our model is capable of generating 1024px images in a single step, achieving quality comparable to state-of-the-art methods.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Diffusion Adversarial Post-Training for One-Step Video Generation | TensorX