TensorX
返回文献探索

Paper · arXiv 2401.15687

Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance

Qingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin, Han Liang, Longwen Zhang, Yingliang Zhang, Jingyi Yu, Lan Xu

24 upvotesJanuary 28, 2024arXiv 预印本
AI 摘要

The proposed Generalized Neural Parametric Facial Asset and Media2Face diffusion model generate high-fidelity and expressive 3D facial animations from speech with rich multi-modality guidance.

variational auto-encoderlatent spaceM2F-D datasetdiffusion model3D facial animationco-speechemotional labelsstyle labelsMedia2Face

Abstract

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism and a lack of lexible conditioning. We address this challenge through a trilogy. We first introduce Generalized Neural Parametric Facial Asset (GNPFA), an efficient variational auto-encoder mapping facial geometry and images to a highly generalized expression latent space, decoupling expressions and identities. Then, we utilize GNPFA to extract high-quality expressions and accurate head poses from a large array of videos. This presents the M2F-D dataset, a large, diverse, and scan-level co-speech 3D facial animation dataset with well-annotated emotional and style labels. Finally, we propose Media2Face, a diffusion model in GNPFA latent space for co-speech facial animation generation, accepting rich multi-modality guidances from audio, text, and image. Extensive experiments demonstrate that our model not only achieves high fidelity in facial animation synthesis but also broadens the scope of expressiveness and style adaptability in 3D facial animation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance | TensorX