TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

31

Platypus: Quick, Cheap, and Powerful Refinement of LLMs

Ariel N. Lee, Cole J. Hunter, Nataniel Ruiz

A fine-tuned and merged family of large language models named Platypus, using LoRA modules and a curated dataset, achieves top performance on the Open LLM Leaderboard with reduced data and compute.

25Large Language ModelsLLMsHF ↗arXiv ↗
41

From Sparse to Soft Mixtures of Experts

Joan Puigcerver, Carlos Riquelme, Basil Mustafa +1 authors

Soft MoE, a differentiable sparse Transformer, stabilizes training, reduces inference cost, and outperforms traditional Transformers and MoE variants in visual recognition.

22sparse mixture of expert architecturesMoEsHF ↗arXiv ↗
48

CausalLM is not optimal for in-context learning

Nan Ding, Tomer Levinboim, Jialin Wu +2 authors

Theoretical analysis shows that prefix language models outperform causal language models in in-context learning by converging to optimal solutions, while causal models exhibit dynamics similar to online gradient descent.

19transformerin-context learningHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号