TensorX
返回文献探索

Paper · arXiv 2401.02038

Understanding LLMs: A Comprehensive Overview from Training to Inference

Yiheng Liu, Hao He, Tianle Han, Xu Zhang, Mengyuan Liu, Jiaming Tian, Yutong Zhang, Jiaqi Wang, Xiaohui Gao, Tianyang Zhong, Yi Pan, Shaochen Xu, Zihao Wu, Zhengliang Liu, Xin Zhang, Shu Zhang, Xintao Hu, Tuo Zhang, Ning Qiang, Tianming Liu, Bao Ge

66 upvotesJanuary 4, 2024arXiv 预印本
AI 摘要

The paper reviews techniques for cost-efficient training and deployment of large language models, covering aspects like data preprocessing, parallel training, model fine-tuning, and inference optimizations including model compression and memory scheduling.

Large Language Modelspre-training tasksparallel trainingmodel fine-tuningmodel compressionparallel computationmemory schedulingstructural optimization

Abstract

The introduction of ChatGPT has led to a significant increase in the utilization of Large Language Models (LLMs) for addressing downstream tasks. There's an increasing focus on cost-efficient training and deployment within this context. Low-cost training and deployment of LLMs represent the future development trend. This paper reviews the evolution of large language model training techniques and inference deployment technologies aligned with this emerging trend. The discussion on training includes various aspects, including data preprocessing, training architecture, pre-training tasks, parallel training, and relevant content related to model fine-tuning. On the inference side, the paper covers topics such as model compression, parallel computation, memory scheduling, and structural optimization. It also explores LLMs' utilization and provides insights into their future development.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Understanding LLMs: A Comprehensive Overview from Training to Inference | TensorX