TensorX
返回文献探索

Paper · arXiv 2312.09767

DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models

Yifeng Ma, Shiwei Zhang, Jiayu Wang, Xiang Wang, Yingya Zhang, Zhidong Deng

27 upvotesDecember 15, 2023arXiv 预印本
AI 摘要

DreamTalk enables expressive talking head generation using diffusion models with specialized components for denoising, lip-sync, and style prediction.

denoising networkstyle-aware lip expertstyle predictordiffusion modelsaudio-driven face motionslip-synctarget expression predictionphoto-realistic talking facesdiverse speaking stylesaccurate lip motions

Abstract

Diffusion models have shown remarkable success in a variety of downstream generative tasks, yet remain under-explored in the important and challenging expressive talking head generation. In this work, we propose a DreamTalk framework to fulfill this gap, which employs meticulous design to unlock the potential of diffusion models in generating expressive talking heads. Specifically, DreamTalk consists of three crucial components: a denoising network, a style-aware lip expert, and a style predictor. The diffusion-based denoising network is able to consistently synthesize high-quality audio-driven face motions across diverse expressions. To enhance the expressiveness and accuracy of lip motions, we introduce a style-aware lip expert that can guide lip-sync while being mindful of the speaking styles. To eliminate the need for expression reference video or text, an extra diffusion-based style predictor is utilized to predict the target expression directly from the audio. By this means, DreamTalk can harness powerful diffusion models to generate expressive faces effectively and reduce the reliance on expensive style references. Experimental results demonstrate that DreamTalk is capable of generating photo-realistic talking faces with diverse speaking styles and achieving accurate lip motions, surpassing existing state-of-the-art counterparts.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models | TensorX