TensorX
返回文献探索

Paper · arXiv 2507.13563

A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models

Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Oleg Rogov, Grach Mkrtchian

53 upvotesJuly 17, 2025arXiv 预印本
AI 摘要

Balalaika, a new Russian speech dataset, improves speech synthesis and enhancement through comprehensive annotations and large-scale data.

speech synthesisvowel reductionconsonant devoicingstress patternshomograph ambiguityintonationdatasettextual annotationspunctuationstress markingscomparative evaluations

Abstract

Russian speech synthesis presents distinctive challenges, including vowel reduction, consonant devoicing, variable stress patterns, homograph ambiguity, and unnatural intonation. This paper introduces Balalaika, a novel dataset comprising more than 2,000 hours of studio-quality Russian speech with comprehensive textual annotations, including punctuation and stress markings. Experimental results show that models trained on Balalaika significantly outperform those trained on existing datasets in both speech synthesis and enhancement tasks. We detail the dataset construction pipeline, annotation methodology, and results of comparative evaluations.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号