TensorX
返回文献探索

Paper · arXiv 2412.19638

Xmodel-2 Technical Report

Wang Qun, Liu Yang, Lin Qingquan, Qu Zhijiu, Jiang Ling

27 upvotesDecember 27, 2024arXiv 预印本
AI 摘要

Xmodel-2, a 1.2-billion-parameter large language model, utilizes a unified set of hyperparameters across different scales and the WSD learning rate scheduler to achieve state-of-the-art performance in complex reasoning tasks with efficient training.

large language modelreasoning tasksunified set of hyperparametersWSD learning rate schedulerMiniCPMpretrainedagent-based tasks

Abstract

Xmodel-2 is a 1.2-billion-parameter large language model designed specifically for reasoning tasks. Its architecture enables different model scales to share a unified set of hyperparameters, allowing for extensive experimentation on smaller models and seamless transfer of optimal configurations to larger models. To maximize training efficiency and stability, Xmodel-2 employs the WSD learning rate scheduler from MiniCPM. Pretrained on 1.5 trillion tokens from diverse sources, Xmodel-2 achieves state-of-the-art performance in complex reasoning and agent-based tasks, while maintaining low training costs. These results highlight the potential of efficient model design and training strategies in advancing reasoning capabilities. Model checkpoints and code are publicly available on GitHub at https://github.com/XiaoduoAILab/Xmodel-2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号