TensorX
返回文献探索

Paper · arXiv 2307.11795

Prompting Large Language Models with Speech Recognition Abilities

Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, Christian Fuegen, Mike Seltzer

18 upvotesJuly 21, 2023arXiv 预印本
AI 摘要

Attaching a conformer audio encoder to LLaMA-7B allows it to perform multilingual speech recognition, outperforming monolingual models and maintaining its original text capabilities, even with limited training and longer strides.

large language modelsspeech recognitionabstractive summarizationopen-ended question answeringautomatic speech recognitionMLSconformer encoderLLaMA-7Bmultilingual ASRablation studiesaudio encoderaudio encoder striding

Abstract

Large language models have proven themselves highly flexible, able to solve a wide range of generative tasks, such as abstractive summarization and open-ended question answering. In this paper we extend the capabilities of LLMs by directly attaching a small audio encoder allowing it to perform speech recognition. By directly prepending a sequence of audial embeddings to the text token embeddings, the LLM can be converted to an automatic speech recognition (ASR) system, and be used in the exact same manner as its textual counterpart. Experiments on Multilingual LibriSpeech (MLS) show that incorporating a conformer encoder into the open sourced LLaMA-7B allows it to outperform monolingual baselines by 18% and perform multilingual speech recognition despite LLaMA being trained overwhelmingly on English text. Furthermore, we perform ablation studies to investigate whether the LLM can be completely frozen during training to maintain its original capabilities, scaling up the audio encoder, and increasing the audio encoder striding to generate fewer embeddings. The results from these studies show that multilingual ASR is possible even when the LLM is frozen or when strides of almost 1 second are used in the audio encoder opening up the possibility for LLMs to operate on long-form audio.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Prompting Large Language Models with Speech Recognition Abilities | TensorX