TensorX
返回文献探索

Paper · arXiv 2405.10626

Dynamic data sampler for cross-language transfer learning in large language models

Yudong Li, Yuhao Feng, Wen Zhou, Zhe Zhao, Linlin Shen, Cheng Hou, Xianxu Hou

6 upvotesMay 17, 2024arXiv 预印本
AI 摘要

ChatFlow, a cross-language transfer-based LLM using Chinese, English, and parallel corpus, enhances training efficiency and performance for Chinese language models leveraging LLaMA2.

Large Language Modelsnatural language processingLLMsNLPChatFlowcross-language transferLLaMA2 modelunsupervised pre-trainingsupervised fine-tuningmodel convergence

Abstract

Large Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for languages other than English poses significant challenges, due to the difficulty in acquiring large-scale corpus and the requisite computing resources. In this paper, we propose ChatFlow, a cross-language transfer-based LLM, to address these challenges and train large Chinese language models in a cost-effective manner. We employ a mix of Chinese, English, and parallel corpus to continuously train the LLaMA2 model, aiming to align cross-language representations and facilitate the knowledge transfer specifically to the Chinese language model. In addition, we use a dynamic data sampler to progressively transition the model from unsupervised pre-training to supervised fine-tuning. Experimental results demonstrate that our approach accelerates model convergence and achieves superior performance. We evaluate ChatFlow on popular Chinese and English benchmarks, the results indicate that it outperforms other Chinese models post-trained on LLaMA-2-7B.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Dynamic data sampler for cross-language transfer learning in large language models | TensorX