TensorX
返回文献探索

Paper · arXiv 2408.03900

Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

Beomseok Lee, Ioan Calapodescu, Marco Gaido, Matteo Negri, Laurent Besacier

10 upvotesAugust 7, 2024arXiv 预印本
AI 摘要

A multilingual spoken language understanding dataset developed for evaluating foundation models across languages and tasks, offering scenarios for zero-shot, few-shot, and full fine-tuning.

spoken language understandingmultilingual datasetintent predictionslot-fillingfoundation modelscascaded architecturesend-to-end architectureszero-shotfew-shotfull fine-tunespeech transcriptionlanguage identificationspeech translation

Abstract

We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. Our extension is prompted by the scarcity of massively multilingual SLU datasets and the growing need for versatile speech datasets to assess foundation models (LLMs, speech encoders) across languages and tasks. We provide a multimodal, multitask, multilingual dataset and report SLU baselines using both cascaded and end-to-end architectures in various training scenarios (zero-shot, few-shot, and full fine-tune). Furthermore, we demonstrate the suitability of Speech-MASSIVE for benchmarking other tasks such as speech transcription, language identification, and speech translation. The dataset, models, and code are publicly available at: https://github.com/hlt-mt/Speech-MASSIVE

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond | TensorX