TensorX
返回文献探索

Paper · arXiv 2601.18184

VIBEVOICE-ASR Technical Report

Zhiliang Peng, Jianwei Yu, Yaoyao Chang, Zilong Wang, Li Dong, Yingbo Hao, Yujie Tu, Chenyu Yang, Wenhui Wang, Songchen Xu, Yutao Sun, Hangbo Bao, Weijiang Xu, Yi Zhu, Zehua Wang, Ting Song, Yan Xia, Zewen Chi, Shaohan Huang, Liang Wang, Chuang Ding, Shuai Wang, Xie Chen, Furu Wei

25 upvotesMarch 14, 2026arXiv 预印本
AI 摘要

VibeVoice-ASR is a unified end-to-end speech understanding framework that processes long-form audio without chunking, supports multiple languages and code-switching, and uses prompt-based context injection for improved domain-specific accuracy.

speech understanding frameworkVibeVoicecontext fragmentationmulti-speaker complexitylong-form audiosingle-pass processingAutomatic Speech RecognitionSpeaker DiarizationTimestampingend-to-end generationcode-switchingprompt-based context injection

Abstract

This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g., meetings, podcasts) that remain despite recent advancements in short-form speech recognition. Unlike traditional pipelined approaches that rely on audio chunking, VibeVoice-ASRsupports single-pass processing for up to 60 minutes of audio. It unifies Automatic Speech Recognition, Speaker Diarization, and Timestamping into a single end-to-end generation task. In addition, VibeVoice-ASR supports over 50 languages, requires no explicit language setting, and natively handles code-switching within and across utterances. Furthermore, we introduce a prompt-based context injection mechanism that allows users to supply customized conetxt, significantly improving accuracy on domain-specific terminology and polyphonic character disambiguation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号