TensorX
返回文献探索

Paper · arXiv 2412.17743

YuLan-Mini: An Open Data-efficient Language Model

Yiwen Hu, Huatong Song, Jia Deng, Jiapeng Wang, Jie Chen, Kun Zhou, Yutao Zhu, Jinhao Jiang, Zican Dong, Wayne Xin Zhao, Ji-Rong Wen

67 upvotesDecember 23, 2024arXiv 预印本
AI 摘要

YuLan-Mini, a 2.42B parameter base model, achieves top-tier performance with efficient pre-training techniques, including a data pipeline with cleaning and scheduling, robust optimization, and annealing with targeted data selection.

data pipelinedata cleaningdata schedule strategiesrobust optimizationannealing approachtargeted data selectionlong context training

Abstract

Effective pre-training of large language models (LLMs) has been challenging due to the immense resource demands and the complexity of the technical processes involved. This paper presents a detailed technical report on YuLan-Mini, a highly capable base model with 2.42B parameters that achieves top-tier performance among models of similar parameter scale. Our pre-training approach focuses on enhancing training efficacy through three key technical contributions: an elaborate data pipeline combines data cleaning with data schedule strategies, a robust optimization method to mitigate training instability, and an effective annealing approach that incorporates targeted data selection and long context training. Remarkably, YuLan-Mini, trained on 1.08T tokens, achieves performance comparable to industry-leading models that require significantly more data. To facilitate reproduction, we release the full details of the data composition for each training phase. Project details can be accessed at the following link: https://github.com/RUC-GSAI/YuLan-Mini.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
YuLan-Mini: An Open Data-efficient Language Model | TensorX