TensorX
返回文献探索

Paper · arXiv 2502.12900

Soundwave: Less is More for Speech-Text Alignment in LLMs

Yuhao Zhang, Zhiheng Liu, Fan Bu, Ruiyu Zhang, Benyou Wang, Haizhou Li

85 upvotesFebruary 18, 2025arXiv 预印本
AI 摘要

Soundwave addresses the representation space gap and sequence length inconsistency in end-to-end speech large language models using an efficient training strategy and novel architecture, outperforming Qwen2-Audio with significantly less data.

large language modelslarge-scale annotated datadata-efficient trainingrepresentation space gapsequence length inconsistencysoundwaveQwen2-Audiospeech translationAIR-Bench speech tasks

Abstract

Existing end-to-end speech large language models (LLMs) usually rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth. We focus on two fundamental problems between speech and text: the representation space gap and sequence length inconsistency. We propose Soundwave, which utilizes an efficient training strategy and a novel architecture to address these issues. Results show that Soundwave outperforms the advanced Qwen2-Audio in speech translation and AIR-Bench speech tasks, using only one-fiftieth of the training data. Further analysis shows that Soundwave still retains its intelligence during conversation. The project is available at https://github.com/FreedomIntelligence/Soundwave.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号