AudioPaLM: A Large Language Model That Can Speak and Listen
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen +27 authors
A unified multimodal language model combining text and speech capabilities outperforms existing systems in speech translation and zero-shot speech-to-text translation.