TensorX
返回文献探索

Paper · arXiv 2508.08248

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

Shuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han, Zhen Xing, Qi Dai, Chong Luo, Zuxuan Wu, Yu-Gang Jiang

27 upvotesAugust 11, 2025arXiv 预印本
AI 摘要

StableAvatar, an end-to-end video diffusion transformer, synthesizes infinite-length high-quality audio-driven avatar videos with natural synchronization and identity consistency using a Time-step-aware Audio Adapter and Audio Native Guidance Mechanism.

diffusion modelsaudio-driven avatar video generationend-to-end video diffusion transformerinfinite-length video generationreference imageaudio embeddingscross-attentiondiffusion backboneslatent distribution error accumulationTime-step-aware Audio AdapterAudio Native Guidance MechanismDynamic Weighted Sliding-window Strategy

Abstract

Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion transformer that synthesizes infinite-length high-quality videos without post-processing. Conditioned on a reference image and audio, StableAvatar integrates tailored training and inference modules to enable infinite-length video generation. We observe that the main reason preventing existing models from generating long videos lies in their audio modeling. They typically rely on third-party off-the-shelf extractors to obtain audio embeddings, which are then directly injected into the diffusion model via cross-attention. Since current diffusion backbones lack any audio-related priors, this approach causes severe latent distribution error accumulation across video clips, leading the latent distribution of subsequent segments to drift away from the optimal distribution gradually. To address this, StableAvatar introduces a novel Time-step-aware Audio Adapter that prevents error accumulation via time-step-aware modulation. During inference, we propose a novel Audio Native Guidance Mechanism to further enhance the audio synchronization by leveraging the diffusion's own evolving joint audio-latent prediction as a dynamic guidance signal. To enhance the smoothness of the infinite-length videos, we introduce a Dynamic Weighted Sliding-window Strategy that fuses latent over time. Experiments on benchmarks show the effectiveness of StableAvatar both qualitatively and quantitatively.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号