TensorX
返回文献探索

Paper · arXiv 2607.10371

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov, Georgii Gospodinov, Dmitrii Bolotov, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin

34 upvotesJuly 11, 2026arXiv 预印本
AI 摘要

A Conformer encoder pre-trained with HuBERT-style objectives and cluster-level balancing improves ASR for underrepresented Central Asian languages.

multilingual ASRConformer encoderHuBERT-style objectivecluster-level data balancingdomain-aware samplingWhisper Large v3Omnilingual-1B

Abstract

Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek). We present GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
GigaAM Multilingual: Foundation Model for Underrepresented Languages | TensorX