TensorX
返回文献探索

Paper · arXiv 2401.12246

Orion-14B: Open-source Multilingual Large Language Models

Du Chen, Yi Huang, Xiaopu Li, Yongqiang Li, Yongqiang Liu, Haihui Pan, Leichao Xu, Dacheng Zhang, Zhipeng Zhang, Kun Han

14 upvotesJanuary 20, 2024arXiv 预印本
AI 摘要

A multilingual large language model with 14 billion parameters achieves state-of-the-art performance across various tasks through data scheduling and fine-tuning.

multilingual large language modelsdata schedulingconversational applicationsstate-of-the-art performance

Abstract

In this study, we introduce Orion-14B, a collection of multilingual large language models with 14 billion parameters. We utilize a data scheduling approach to train a foundational model on a diverse corpus of 2.5 trillion tokens, sourced from texts in English, Chinese, Japanese, Korean, and other languages. Additionally, we fine-tuned a series of models tailored for conversational applications and other specific use cases. Our evaluation results demonstrate that Orion-14B achieves state-of-the-art performance across a broad spectrum of tasks. We make the Orion-14B model family and its associated code publicly accessible https://github.com/OrionStarAI/Orion, aiming to inspire future research and practical applications in the field.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Orion-14B: Open-source Multilingual Large Language Models | TensorX