TensorX
返回文献探索

Paper · arXiv 2404.16645

Tele-FLM Technical Report

Xiang Li, Yiqun Yao, Xin Jiang, Xuezhi Fang, Chao Wang, Xinzhang Liu, Zihan Wang, Yu Zhao, Xin Wang, Yuyao Huang, Shuangyong Song, Yongxiang Li, Zheng Zhang, Bo Zhao, Aixin Sun, Yequan Wang, Zhongjiang He, Zhongyuan Wang, Xuelong Li, Tiejun Huang

18 upvotesApril 25, 2024arXiv 预印本
AI 摘要

An open-sourced 52B multilingual large language model, Tele-FLM, features efficient pre-training and superior language capabilities, and is comparable to larger models in terms of performance metrics.

LLMslarge language modelspre-training paradigmfactual judgment capabilitiesBPBmultilingual language modelingEnglish foundation modelChinese foundation modelFLOPsLlama2-70BDeepSeek-67B

Abstract

Large language models (LLMs) have showcased profound capabilities in language understanding and generation, facilitating a wide array of applications. However, there is a notable paucity of detailed, open-sourced methodologies on efficiently scaling LLMs beyond 50 billion parameters with minimum trial-and-error cost and computational resources. In this report, we introduce Tele-FLM (aka FLM-2), a 52B open-sourced multilingual large language model that features a stable, efficient pre-training paradigm and enhanced factual judgment capabilities. Tele-FLM demonstrates superior multilingual language modeling abilities, measured by BPB on textual corpus. Besides, in both English and Chinese foundation model evaluation, it is comparable to strong open-sourced models that involve larger pre-training FLOPs, such as Llama2-70B and DeepSeek-67B. In addition to the model weights, we share the core designs, engineering practices, and training details, which we expect to benefit both the academic and industrial communities.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Tele-FLM Technical Report | TensorX