TensorX
返回文献探索

Paper · arXiv 2401.01885

From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations

Evonne Ng, Javier Romero, Timur Bagautdinov, Shaojie Bai, Trevor Darrell, Angjoo Kanazawa, Alexander Richard

28 upvotesJanuary 3, 2024arXiv 预印本
AI 摘要

A framework combines vector quantization and diffusion models to generate diverse, dynamic, and photorealistic avatars responding to speech in a conversational setting using a novel multi-view dataset.

vector quantizationdiffusionphotorealistic avatarsgestural motionfacebodyhandsmulti-view conversational datasetperceptual evaluation

Abstract

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural motion for an individual, including face, body, and hands. The key behind our method is in combining the benefits of sample diversity from vector quantization with the high-frequency details obtained through diffusion to generate more dynamic, expressive motion. We visualize the generated motion using highly photorealistic avatars that can express crucial nuances in gestures (e.g. sneers and smirks). To facilitate this line of research, we introduce a first-of-its-kind multi-view conversational dataset that allows for photorealistic reconstruction. Experiments show our model generates appropriate and diverse gestures, outperforming both diffusion- and VQ-only methods. Furthermore, our perceptual evaluation highlights the importance of photorealism (vs. meshes) in accurately assessing subtle motion details in conversational gestures. Code and dataset available online.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations | TensorX