TensorX
返回文献探索

Paper · arXiv 2401.13303

MaLA-500: Massive Language Adaptation of Large Language Models

Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann, André F. T. Martins, Hinrich Schütze

11 upvotesJanuary 24, 2024arXiv 预印本
AI 摘要

MaLA-500, a large language model covering 534 languages, achieves state-of-the-art in-context learning by extending the LLaMA 2 vocabulary with Glot500-c.

MaLA-500large language modelsin-context learningLLaMA 2vocabulary extensionGlot500-c

Abstract

Large language models have advanced the state of the art in natural language processing. However, their predominant design for English or a limited set of languages creates a substantial gap in their effectiveness for low-resource languages. To bridge this gap, we introduce MaLA-500, a novel large language model designed to cover an extensive range of 534 languages. To train MaLA-500, we employ vocabulary extension and continued pretraining on LLaMA 2 with Glot500-c. Our experiments on SIB-200 show that MaLA-500 achieves state-of-the-art in-context learning results. We release MaLA-500 at https://huggingface.co/MaLA-LM

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MaLA-500: Massive Language Adaptation of Large Language Models | TensorX