TensorX
返回文献探索

Paper · arXiv 2510.24717

Uniform Discrete Diffusion with Metric Path for Video Generation

Haoge Deng, Ting Pan, Fan Zhang, Yang Liu, Zhuoyan Luo, Yufeng Cui, Wenxuan Wang, Chunhua Shen, Shiguang Shan, Zhaoxiang Zhang, Xinlong Wang

43 upvotesOctober 28, 2025arXiv 预印本
AI 摘要

URSA, a discrete generative model, bridges the gap with continuous approaches in video generation by using iterative refinement, linearized metric paths, and resolution-dependent timestep shifting, achieving performance comparable to state-of-the-art continuous methods.

discrete generative modelingUniform discRete diffuSion with metric pAth (URSA)iterative global refinementdiscrete spatiotemporal tokensLinearized Metric PathResolution-dependent Timestep Shifting mechanismasynchronous temporal fine-tuningvideo generationimage synthesislong-duration video generationinterpolationimage-to-video generation

Abstract

Continuous-space video generation has advanced rapidly, while discrete approaches lag behind due to error accumulation and long-context inconsistency. In this work, we revisit discrete generative modeling and present Uniform discRete diffuSion with metric pAth (URSA), a simple yet powerful framework that bridges the gap with continuous approaches for the scalable video generation. At its core, URSA formulates the video generation task as an iterative global refinement of discrete spatiotemporal tokens. It integrates two key designs: a Linearized Metric Path and a Resolution-dependent Timestep Shifting mechanism. These designs enable URSA to scale efficiently to high-resolution image synthesis and long-duration video generation, while requiring significantly fewer inference steps. Additionally, we introduce an asynchronous temporal fine-tuning strategy that unifies versatile tasks within a single model, including interpolation and image-to-video generation. Extensive experiments on challenging video and image generation benchmarks demonstrate that URSA consistently outperforms existing discrete methods and achieves performance comparable to state-of-the-art continuous diffusion methods. Code and models are available at https://github.com/baaivision/URSA

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Uniform Discrete Diffusion with Metric Path for Video Generation | TensorX