TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

511

Transformers are Multi-State RNNs

Matanel Oren, Michael Hassid, Yossi Adi +1 authors

Decoder-only transformers can be conceptualized as finite multi-state RNNs, and a new cache compression technique, TOVA, significantly reduces their computational cost while maintaining high performance.

39transformersrecurrent neural networksHF ↗arXiv ↗
513

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Philip Anastassiou, Jiawei Chen, Jitong Chen +43 authors

Seed-TTS is a family of large-scale TTS models that generate high-quality speech with in-context learning, superior controllability, and a non-autoregressive variant using diffusion-based architecture that does not rely on pre-estimated phoneme durations.

39autoregressive text-to-speechspeech generationHF ↗arXiv ↗
516

JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Yikang Shen, Zhen Guo, Tianle Cai +1 authors

JetMoE-8B, a cost-effective large language model with 8 billion parameters, achieves impressive performance using a Sparsely-gated Mixture-of-Experts architecture, demonstrating efficient use of resources and computational savings.

38Sparsely-gated Mixture-of-ExpertsSMoEHF ↗arXiv ↗
517

SUTRA: Scalable Multilingual Language Model Architecture

Abhijit Bendale, Michael Sapienza, Steven Ripplinger +3 authors

SUTRA, a multilingual Large Language Model architecture, achieves superior performance on multilingual tasks by decoupling conceptual understanding from language-specific processing using a Mixture of Experts framework.

38Multilingual Large Language ModelMixture of Experts frameworkHF ↗arXiv ↗
529

Scalable Pre-training of Large Autoregressive Image Models

Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai +5 authors

Autoregressive pre-training for vision models (AIM) scales similarly to LLMs, showing improved performance with more data and parameters, and does not exhibit performance saturation.

38autoregressive objectivevision modelsHF ↗arXiv ↗
535

Long-context LLMs Struggle with Long In-context Learning

Tianle Li, Ge Zhang, Quy Duc Do +2 authors

LIConBench evaluates long-context LLMs on extreme-label classification tasks with sequences up to 50K tokens, highlighting performance dips beyond 20K tokens and favoring of recent labels.

37Large Language Modelslong in-context learningHF ↗arXiv ↗
540

jina-embeddings-v3: Multilingual Embeddings With Task LoRA

Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram +9 authors

jina-embeddings-v3, a large-scale text embedding model, achieves state-of-the-art performance in multilingual and long-context retrieval tasks using Low-Rank Adaptation and Matryoshka Representation Learning.

37Low-Rank AdaptationLoRA adaptersHF ↗arXiv ↗
18 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号