TensorX
返回文献探索

Paper · arXiv 2508.21290

Efficient Code Embeddings from Code Generation Models

Daria Kryvosheieva, Saba Sturua, Michael Günther, Scott Martens, Han Xiao

21 upvotesAugust 29, 2025arXiv 预印本
AI 摘要

Jina-code-embeddings uses an autoregressive backbone pre-trained on text and code to generate embeddings for code retrieval, question-answering, and similarity identification.

autoregressive backbonelast-token poolingcode embedding model

Abstract

jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages. It makes innovative use of an autoregressive backbone pre-trained on both text and code, generating embeddings via last-token pooling. We outline the training recipe and demonstrate state-of-the-art performance despite the relatively small size of the models, validating this approach to code embedding model construction.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号