TensorX
返回文献探索

Paper · arXiv 2404.01856

Poro 34B and the Blessing of Multilinguality

Risto Luukkonen, Jonathan Burdge, Elaine Zosa, Aarne Talman, Ville Komulainen, Väinö Hatanpää, Peter Sarlin, Sampo Pyysalo

13 upvotesApril 2, 2024arXiv 预印本
AI 摘要

A multilingual training approach on a large language model improves capabilities for small languages, translation, and generation in multiple languages.

large language modelspretrainingmultilingualitymonolingual modelsparameter modeltrainingtransliterationtranslationgeneration

Abstract

The pretraining of state-of-the-art large language models now requires trillions of words of text, which is orders of magnitude more than available for the vast majority of languages. While including text in more than one language is an obvious way to acquire more pretraining data, multilinguality is often seen as a curse, and most model training efforts continue to focus near-exclusively on individual large languages. We believe that multilinguality can be a blessing and that it should be possible to substantially improve over the capabilities of monolingual models for small languages through multilingual training. In this study, we introduce Poro 34B, a 34 billion parameter model trained for 1 trillion tokens of Finnish, English, and programming languages, and demonstrate that a multilingual training approach can produce a model that not only substantially advances over the capabilities of existing models for Finnish, but also excels in translation and is competitive in its class in generating English and programming languages. We release the model parameters, scripts, and data under open licenses at https://huggingface.co/LumiOpen/Poro-34B.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Poro 34B and the Blessing of Multilinguality | TensorX