TensorX
返回文献探索

Paper · arXiv 2403.17694

AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Huawei Wei, Zejun Yang, Zhisheng Wang

11 upvotesMarch 26, 2024arXiv 预印本
AI 摘要

AniPortrait generates high-quality portrait animations from audio and a reference image using 3D intermediate representations, 2D facial landmarks, and a diffusion model with a motion module.

3D intermediate representations2D facial landmarksdiffusion modelmotion module

Abstract

In this study, we propose AniPortrait, a novel framework for generating high-quality animation driven by audio and a reference portrait image. Our methodology is divided into two stages. Initially, we extract 3D intermediate representations from audio and project them into a sequence of 2D facial landmarks. Subsequently, we employ a robust diffusion model, coupled with a motion module, to convert the landmark sequence into photorealistic and temporally consistent portrait animation. Experimental results demonstrate the superiority of AniPortrait in terms of facial naturalness, pose diversity, and visual quality, thereby offering an enhanced perceptual experience. Moreover, our methodology exhibits considerable potential in terms of flexibility and controllability, which can be effectively applied in areas such as facial motion editing or face reenactment. We release code and model weights at https://github.com/scutzzj/AniPortrait

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation | TensorX