TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

31

TransformerFAM: Feedback attention is working memory

Dongseong Hwang, Weiran Wang, Zhuoyuan Huo +2 authors

Feedback Attention Memory (FAM) enhances Transformer architecture by enabling long-context processing without additional weights, significantly improving performance on large sequences across various model sizes.

43TransformersFeedback Attention MemoryHF ↗arXiv ↗
38

JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Yikang Shen, Zhen Guo, Tianle Cai +1 authors

JetMoE-8B, a cost-effective large language model with 8 billion parameters, achieves impressive performance using a Sparsely-gated Mixture-of-Experts architecture, demonstrating efficient use of resources and computational savings.

38Sparsely-gated Mixture-of-ExpertsSMoEHF ↗arXiv ↗
39

Long-context LLMs Struggle with Long In-context Learning

Tianle Li, Ge Zhang, Quy Duc Do +2 authors

LIConBench evaluates long-context LLMs on extreme-label classification tasks with sequences up to 50K tokens, highlighting performance dips beyond 20K tokens and favoring of recent labels.

37Large Language Modelslong in-context learningHF ↗arXiv ↗
40

Pre-training Small Base LMs with Fewer Tokens

Sunny Sanyal, Sujay Sanghavi, Alexandros G. Dimakis

Inheritune leverages transformer blocks from large language models and minimal data to create smaller, efficient base models with performance competitive to larger models.

36transformer blocksInherituneHF ↗arXiv ↗
42

Pegasus-v1 Technical Report

Raehyuk Jung, Hyojun Go, Jaehyuk Yi +41 authors

Pegasus-1 is a multimodal language model designed for video content comprehension, handling spatiotemporal information and demonstrated in benchmarks for video conversation, zero-shot video question answering, and video summarization.

34multimodal language modelvideo content understandingHF ↗arXiv ↗
44

FlowMind: Automatic Workflow Generation with LLMs

Zhen Zeng, William Watson, Nicole Cho +4 authors

FlowMind uses Large Language Models with a generic prompt recipe to generate automatic workflows, addressing spontaneous tasks and ensuring data integrity, and it is evaluated using a new financial dataset NCEN-QA.

34Large Language ModelsGenerative Pretrained TransformerHF ↗arXiv ↗
47

FlashSpeech: Efficient Zero-Shot Speech Synthesis

Zhen Ye, Zeqian Ju, Haohe Liu +10 authors

FlashSpeech, using latent consistency model and adversarial consistency training, achieves fast and high-quality zero-shot speech synthesis with enhanced prosody generation.

32latent consistency modeladversarial consistency trainingHF ↗arXiv ↗
48

Best Practices and Lessons Learned on Synthetic Data for Language Models

Ruibo Liu, Jerry Wei, Fangyu Liu +8 authors

The success of AI models relies on the availability of large, diverse, and high-quality datasets, which can be challenging to obtain due to data scarcity, privacy concerns, and high costs. Synthetic data has emerged as a promising solution by generating artificial data that mimics real-world patterns. This paper provides an overview of synthetic data research, discussing its applications, challenges, and future directions. We present empirical evidence from prior art to demonstrate its effectiveness and highlight the importance of ensuring its factuality, fidelity, and unbiasedness. We emphasize the need for responsible use of synthetic data to build more powerful, inclusive, and trustworthy language models.

32HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号