TensorX
返回文献探索

Paper · arXiv 2407.13739

Scaling Granite Code Models to 128K Context

Matt Stallone, Vaibhav Saxena, Leonid Karlinsky, Bridget McGinn, Tim Bula, Mayank Mishra, Adriana Meza Soria, Gaoyuan Zhang, Aditya Prasad, Yikang Shen, Saptha Surendran, Shanmukha Guttula, Hima Patel, Parameswaran Selvam, Xuan-Hong Dang, Yan Koyfman, Atin Sood, Rogerio Feris, Nirmit Desai, David D. Cox, Ruchir Puri, Rameswar Panda

21 upvotesJuly 18, 2024arXiv 预印本
AI 摘要

Granite code models with extended context windows up to 128K tokens are achieved through lightweight continual pretraining and finetuning with increased RoPE frequencies and length-upsampled data, showing improvements in long-context tasks without degrading performance on standard benchmarks.

Granite code modelsRoPErepository-level file packinglength-upsampled long-context datainstruction-tuned modelsHumanEvallong-context taskscode completion benchmarks

Abstract

This paper introduces long-context Granite code models that support effective context windows of up to 128K tokens. Our solution for scaling context length of Granite 3B/8B code models from 2K/4K to 128K consists of a light-weight continual pretraining by gradually increasing its RoPE base frequency with repository-level file packing and length-upsampled long-context data. Additionally, we also release instruction-tuned models with long-context support which are derived by further finetuning the long context base models on a mix of permissively licensed short and long-context instruction-response pairs. While comparing to the original short-context Granite code models, our long-context models achieve significant improvements on long-context tasks without any noticeable performance degradation on regular code completion benchmarks (e.g., HumanEval). We release all our long-context Granite code models under an Apache 2.0 license for both research and commercial use.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Scaling Granite Code Models to 128K Context | TensorX