TensorX
返回文献探索

Paper · arXiv 2408.12885

T3M: Text Guided 3D Human Motion Synthesis from Speech

Wenshuo Peng, Kaipeng Zhang, Sai Qian Zhang

13 upvotesAugust 23, 2024arXiv 预印本
AI 摘要

A novel text-guided 3D motion synthesis method called T3M provides precise control over motion generation, offering improved quality and diversity compared to existing speech-driven approaches.

3D motion synthesisspeech-driventextual inputstate-of-the-art methods

Abstract

Speech-driven 3D motion synthesis seeks to create lifelike animations based on human speech, with potential uses in virtual reality, gaming, and the film production. Existing approaches reply solely on speech audio for motion generation, leading to inaccurate and inflexible synthesis results. To mitigate this problem, we introduce a novel text-guided 3D human motion synthesis method, termed T3M. Unlike traditional approaches, T3M allows precise control over motion synthesis via textual input, enhancing the degree of diversity and user customization. The experiment results demonstrate that T3M can greatly outperform the state-of-the-art methods in both quantitative metrics and qualitative evaluations. We have publicly released our code at https://github.com/Gloria2tt/T3M.git{https://github.com/Gloria2tt/T3M.git}

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号