TensorX
返回文献探索

Paper · arXiv 2503.05500

EuroBERT: Scaling Multilingual Encoders for European Languages

Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, Pierre Colombo

81 upvotesMarch 7, 2025arXiv 预印本
AI 摘要

EuroBERT, a family of multilingual encoders covering European and global languages, outperforms existing models across various tasks and supports long sequences, surpassing traditional bidirectional encoders.

bidirectional encoder modelsgenerative decoder-only modelsmultilingual encodersEuroBERTmultilingual capabilitiestoken sequences

Abstract

General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by advances in generative decoder-only models. However, many innovations driving this progress are not inherently tied to decoders. In this paper, we revisit the development of multilingual encoders through the lens of these advances, and introduce EuroBERT, a family of multilingual encoders covering European and widely spoken global languages. Our models outperform existing alternatives across a diverse range of tasks, spanning multilingual capabilities, mathematics, and coding, and natively supporting sequences of up to 8,192 tokens. We also examine the design decisions behind EuroBERT, offering insights into our dataset composition and training pipeline. We publicly release the EuroBERT models, including intermediate training checkpoints, together with our training framework.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
EuroBERT: Scaling Multilingual Encoders for European Languages | TensorX