TensorX
返回文献探索

Paper · arXiv 2308.15930

LLaSM: Large Language and Speech Model

Yu Shu, Siwei Dong, Guangyao Chen, Wenhao Huang, Ruihua Zhang, Daochen Shi, Qiqi Xiang, Yemin Shi

34 upvotesAugust 30, 2023arXiv 预印本
AI 摘要

LLaSM, an end-to-end trained large multi-modal speech-language model, enhances human interaction with AI by following speech-and-language instructions using cross-modal conversational abilities.

multi-modal large language modelsvision-language multi-modal modelsLLaSMcross-modal conversational abilitiesSpeech Instruction Following datasetLLaSM-Audio-Instructions

Abstract

Multi-modal large language models have garnered significant interest recently. Though, most of the works focus on vision-language multi-modal models providing strong capabilities in following vision-and-language instructions. However, we claim that speech is also an important modality through which humans interact with the world. Hence, it is crucial for a general-purpose assistant to be able to follow multi-modal speech-and-language instructions. In this work, we propose Large Language and Speech Model (LLaSM). LLaSM is an end-to-end trained large multi-modal speech-language model with cross-modal conversational abilities, capable of following speech-and-language instructions. Our early experiments show that LLaSM demonstrates a more convenient and natural way for humans to interact with artificial intelligence. Specifically, we also release a large Speech Instruction Following dataset LLaSM-Audio-Instructions. Code and demo are available at https://github.com/LinkSoul-AI/LLaSM and https://huggingface.co/spaces/LinkSoul/LLaSM. The LLaSM-Audio-Instructions dataset is available at https://huggingface.co/datasets/LinkSoul/LLaSM-Audio-Instructions.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLaSM: Large Language and Speech Model | TensorX