TensorX
返回文献探索

Paper · arXiv 2508.05305

SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens

Nikita Dragunov, Temurbek Rahmatullaev, Elizaveta Goncharova, Andrey Kuznetsov, Anton Razzhigaev

49 upvotesAugust 7, 2025arXiv 预印本
AI 摘要

SONAR-LLM, a decoder-only transformer using token-level cross-entropy in the SONAR embedding space, achieves competitive text generation quality without diffusion sampling.

Large Concept ModelLCMSONAR-LLMdecoder-only transformerSONAR embedding spacetoken-level cross-entropydiffusion objectiveslikelihood-based training signal

Abstract

The recently proposed Large Concept Model (LCM) generates text by predicting a sequence of sentence-level embeddings and training with either mean-squared error or diffusion objectives. We present SONAR-LLM, a decoder-only transformer that "thinks" in the same continuous SONAR embedding space, yet is supervised through token-level cross-entropy propagated via the frozen SONAR decoder. This hybrid objective retains the semantic abstraction of LCM while eliminating its diffusion sampler and restoring a likelihood-based training signal. Across model sizes from 39M to 1.3B parameters, SONAR-LLM attains competitive generation quality. We report scaling trends, ablations, benchmark results, and release the complete training code and all pretrained checkpoints to foster reproducibility and future research.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens | TensorX