TensorX
返回文献探索

Paper · arXiv 2409.16235

EuroLLM: Multilingual Language Models for Europe

Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, André F. T. Martins

29 upvotesSeptember 24, 2024arXiv 预印本
AI 摘要

The EuroLLM project develops multilingual language models for European Union languages, detailing data collection, tokenizer creation, and model performance.

open-weight LLMsEuroLLM projectmultilingual LLMsscaling lawsmultilingual tokenizerdata mixmodeling configurationsEuroLLM-1.7BEuroLLM-1.7B-Instruct

Abstract

The quality of open-weight LLMs has seen significant improvement, yet they remain predominantly focused on English. In this paper, we introduce the EuroLLM project, aimed at developing a suite of open-weight multilingual LLMs capable of understanding and generating text in all official European Union languages, as well as several additional relevant languages. We outline the progress made to date, detailing our data collection and filtering process, the development of scaling laws, the creation of our multilingual tokenizer, and the data mix and modeling configurations. Additionally, we release our initial models: EuroLLM-1.7B and EuroLLM-1.7B-Instruct and report their performance on multilingual general benchmarks and machine translation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
EuroLLM: Multilingual Language Models for Europe | TensorX