TensorX
返回文献探索

Paper · arXiv 2401.02385

TinyLlama: An Open-Source Small Language Model

Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, Wei Lu

96 upvotesJanuary 4, 2024arXiv 预印本
AI 摘要

TinyLlama, a compact 1.1B language model, leverages FlashAttention to achieve high performance in downstream tasks with enhanced computational efficiency.

FlashAttentionLlama 2TinyLlamalanguage model

Abstract

We present TinyLlama, a compact 1.1B language model pretrained on around 1 trillion tokens for approximately 3 epochs. Building on the architecture and tokenizer of Llama 2, TinyLlama leverages various advances contributed by the open-source community (e.g., FlashAttention), achieving better computational efficiency. Despite its relatively small size, TinyLlama demonstrates remarkable performance in a series of downstream tasks. It significantly outperforms existing open-source language models with comparable sizes. Our model checkpoints and code are publicly available on GitHub at https://github.com/jzhang38/TinyLlama.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
TinyLlama: An Open-Source Small Language Model | TensorX