TensorX
返回文献探索

Paper · arXiv 2502.01061

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng, Chao Liang

225 upvotesFebruary 3, 2025arXiv 预印本
AI 摘要

OmniHuman, a Diffusion Transformer-based framework, achieves realistic human video generation by integrating motion-related conditions into its training process, supporting various inputs and driving modalities.

Diffusion Transformermotion-related conditionsmodel architectureinference strategydata-driven motion generationhuman video generationportrait contentshuman-object interactionsbody posesimage stylesaudio-drivenvideo-drivencombined driving signals

Abstract

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propose OmniHuman, a Diffusion Transformer-based framework that scales up data by mixing motion-related conditions into the training phase. To this end, we introduce two training principles for these mixed conditions, along with the corresponding model architecture and inference strategy. These designs enable OmniHuman to fully leverage data-driven motion generation, ultimately achieving highly realistic human video generation. More importantly, OmniHuman supports various portrait contents (face close-up, portrait, half-body, full-body), supports both talking and singing, handles human-object interactions and challenging body poses, and accommodates different image styles. Compared to existing end-to-end audio-driven methods, OmniHuman not only produces more realistic videos, but also offers greater flexibility in inputs. It also supports multiple driving modalities (audio-driven, video-driven and combined driving signals). Video samples are provided on the ttfamily project page (https://omnihuman-lab.github.io)

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号