TensorX
返回文献探索

Paper · arXiv 2312.15166

SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling

Dahyun Kim, Chanjun Park, Sanghoon Kim, Wonsung Lee, Wonho Song, Yunsu Kim, Hyeonwoo Kim, Yungi Kim, Hyeonju Lee, Jihoo Kim, Changbae Ahn, Seonghoon Yang, Sukyung Lee, Hyunbyung Park, Gyoungjin Gim, Mikyoung Cha, Hwalsuk Lee, Sunghun Kim

62 upvotesDecember 23, 2023arXiv 预印本
AI 摘要

A novel technique, depth up-scaling (DUS), efficiently enhances large language models (LLMs) without complex changes, building SOLAR 10.7B that outperforms existing open-source models in various NLP tasks, including instruction-following.

depth up-scalingDUSmixture-of-expertsMoElarge language modelLLMSOLAR 10.7BNLPLlama 2Mistral 7BSOLAR 10.7B-InstructMixtral-8x7BApache 2.0 license

Abstract

We introduce depth up-scaling (DUS), a novel technique to up-scale base LLMs efficiently and effectively in a simple manner. In contrast to mixture-of-experts (MoE), DUS does not require complex changes to train and inference. Using DUS, we build SOLAR 10.7B, a large language model (LLM) with 10.7 billion parameters, demonstrating superior performance in various natural language processing (NLP) tasks. Comparative evaluations show that SOLAR 10.7B outperforms existing open-source pretrained LLMs, such as Llama 2 and Mistral 7B. We additionally present SOLAR 10.7B-Instruct, a variant fine-tuned for instruction-following capabilities, surpassing Mixtral-8x7B. SOLAR 10.7B is publicly available under the Apache 2.0 license, promoting broad access and application in the LLM field.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling | TensorX