TensorX
返回文献探索

Paper · arXiv 2503.00865

Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers

Yiran Zhao, Chaoqun Liu, Yue Deng, Jiahao Ying, Mahani Aljunied, Zhaodonghui Li, Lidong Bing, Hou Pong Chan, Yu Rong, Deli Zhao, Wenxuan Zhang

64 upvotesMarch 2, 2025arXiv 预印本
AI 摘要

Babel is an open multilingual LLM that expands its parameter count through layer extension, covering numerous languages and achieving superior performance in multilingual tasks compared to other open LLMs.

LLMsnatural language processingmultilingual LLMslayer extensionBabel-9BBabel-83Binferencefine-tuningsupervised fine-tuningmultilingual tasks

Abstract

Large language models (LLMs) have revolutionized natural language processing (NLP), yet open-source multilingual LLMs remain scarce, with existing models often limited in language coverage. Such models typically prioritize well-resourced languages, while widely spoken but under-resourced languages are often overlooked. To address this disparity, we introduce Babel, an open multilingual LLM that covers the top 25 languages by number of speakers, supports over 90% of the global population, and includes many languages neglected by other open multilingual LLMs. Unlike traditional continue pretraining approaches, Babel expands its parameter count through a layer extension technique that elevates Babel's performance ceiling. We introduce two variants: Babel-9B, designed for efficient inference and fine-tuning, and Babel-83B, which sets a new standard for open multilingual LLMs. Extensive evaluations on multilingual tasks demonstrate its superior performance compared to open LLMs of comparable size. In addition, using open-source supervised fine-tuning datasets, Babel achieves remarkable performance, with Babel-9B-Chat leading among 10B-sized LLMs and Babel-83B-Chat setting a new standard for multilingual tasks, reaching the same level of commercial models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号