TensorX
返回文献探索

Paper · arXiv 2405.17405

Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer

Ruizhi Shao, Youxin Pang, Zerong Zheng, Jingxiang Sun, Yebin Liu

16 upvotesMay 27, 2024arXiv 预印本
AI 摘要
U-Netsdiffusion transformerscascaded 4D transformer architectureattentionhuman identitycamera parameterstemporal signalsmulti-dimensional datasetmulti-dimensional training strategyGANUNet-based diffusion modelsrealisticcoherentfree-view human videosvirtual realityanimation

Abstract

We present a novel approach for generating high-quality, spatio-temporally coherent human videos from a single image under arbitrary viewpoints. Our framework combines the strengths of U-Nets for accurate condition injection and diffusion transformers for capturing global correlations across viewpoints and time. The core is a cascaded 4D transformer architecture that factorizes attention across views, time, and spatial dimensions, enabling efficient modeling of the 4D space. Precise conditioning is achieved by injecting human identity, camera parameters, and temporal signals into the respective transformers. To train this model, we curate a multi-dimensional dataset spanning images, videos, multi-view data and 3D/4D scans, along with a multi-dimensional training strategy. Our approach overcomes the limitations of previous methods based on GAN or UNet-based diffusion models, which struggle with complex motions and viewpoint changes. Through extensive experiments, we demonstrate our method's ability to synthesize realistic, coherent and free-view human videos, paving the way for advanced multimedia applications in areas such as virtual reality and animation. Our project website is https://human4dit.github.io.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer | TensorX