TensorX
返回文献探索

Paper · arXiv 2402.05672

Multilingual E5 Text Embeddings: A Technical Report

Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei

23 upvotesFebruary 8, 2024arXiv 预印本
AI 摘要

Multilingual E5 text embedding models are released in different sizes and include an instruction-tuned variant, achieving performance comparable to state-of-the-art English-only models.

contrastive pre-trainingfine-tuninginstruction-tunedembedding models

Abstract

This technical report presents the training methodology and evaluation results of the open-source multilingual E5 text embedding models, released in mid-2023. Three embedding models of different sizes (small / base / large) are provided, offering a balance between the inference efficiency and embedding quality. The training procedure adheres to the English E5 model recipe, involving contrastive pre-training on 1 billion multilingual text pairs, followed by fine-tuning on a combination of labeled datasets. Additionally, we introduce a new instruction-tuned embedding model, whose performance is on par with state-of-the-art, English-only models of similar sizes. Information regarding the model release can be found at https://github.com/microsoft/unilm/tree/master/e5 .

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Multilingual E5 Text Embeddings: A Technical Report | TensorX