TensorX
返回文献探索

Paper · arXiv 2305.11778

Cross-Lingual Supervision improves Large Language Models Pre-training

Andrea Schioppa, Xavier Garcia, Orhan Firat

2 upvotesMay 19, 2023arXiv 预印本
AI 摘要

Combining self-supervised and supervised objectives during pre-training of Large Language Models improves their in-context learning by incorporating cross-lingual parallel data.

Large Language Modelsself-supervised language modelingnext token predictionspan corruptionMachine Translation Systemscross-lingual supervisionin-context learningpre-trainingsupervised Machine Translation objective

Abstract

The recent rapid progress in pre-training Large Language Models has relied on using self-supervised language modeling objectives like next token prediction or span corruption. On the other hand, Machine Translation Systems are mostly trained using cross-lingual supervision that requires aligned data between source and target languages. We demonstrate that pre-training Large Language Models on a mixture of a self-supervised Language Modeling objective and the supervised Machine Translation objective, therefore including cross-lingual parallel data during pre-training, yields models with better in-context learning abilities. As pre-training is a very resource-intensive process and a grid search on the best mixing ratio between the two objectives is prohibitively expensive, we propose a simple yet effective strategy to learn it during pre-training.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Cross-Lingual Supervision improves Large Language Models Pre-training | TensorX