VibeVoice Technical Report
Zhiliang Peng, Jianwei Yu, Wenhui Wang +10 authors
VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.