TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 5 – Feb 11, 2024
本周最热150

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Zhihong Shao, Peiyi Wang, Qihao Zhu +6 authors

DeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.

DeepSeekMath 7BDeepSeek-Coder-Base-v1.5 7BMATH benchmarkGroup Relative Policy OptimizationHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Grandmaster-Level Chess Without Search

Anian Ruoss, Grégoire Delétang, Sourabh Medapati +5 authors

A large-scale transformer model trained on a vast dataset of chess games outperforms traditional chess engines and other state-of-the-art models without domain-specific tweaks.

70transformer modelsupervised learningHF ↗arXiv ↗
04

Training-Free Consistent Text-to-Image Generation

Yoad Tewel, Omri Kaduri, Rinon Gal +4 authors

ConsiStory achieves state-of-the-art text-to-image subject consistency without fine-tuning by using shared internal activations, attention blocks, and feature injection.

67text-to-image modelssubject consistencyHF ↗arXiv ↗
05

More Agents Is All You Need

Junyou Li, Qin Zhang, Yangbin Yu +2 authors

A sampling-and-voting method enhances large language models' performance by increasing the number of agents, with effectiveness tied to task difficulty.

59large language modelsLLMSHF ↗arXiv ↗
07

Specialized Language Models with Cheap Inference from Limited Domain Data

David Grangier, Angelos Katharopoulos, Pierre Ablin +1 authors

The study evaluates various machine learning approaches under constrained inference and training budgets, demonstrating that hyper-networks and mixture of experts perform well with large pretraining budgets, while small models trained on importance sampled datasets are better with large specialization budgets.

47large language modelspretraining budgetHF ↗arXiv ↗
15

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Quan Sun, Jinsheng Wang, Qiying Yu +4 authors

EVA-CLIP-18B, a 18-billion parameter open-source CLIP model, achieves exceptional zero-shot performance across multiple image classification benchmarks with a relatively small dataset.

31contrastive language-image pretrainingCLIPHF ↗arXiv ↗
16

An Interactive Agent Foundation Model

Zane Durante, Bidipta Sarkar, Ran Gong +17 authors

The proposed Interactive Agent Foundation Model uses a multi-task agent training paradigm to develop versatile AI agents capable of performing well across diverse domains and tasks.

29multi-task agent training paradigmvisual masked auto-encodersHF ↗arXiv ↗
18

OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Fuzhao Xue, Zian Zheng, Yao Fu +4 authors

OpenMoE, a series of open-sourced MoE-based decoder-only LLMs, demonstrates improved cost-effectiveness over dense LLMs and reveals insights into MoE routing mechanisms, including context-independent specialization, early routing learning, and drop-towards-the-end issues.

28Mixture-of-Experts (MoE)decoder-onlyHF ↗arXiv ↗
25

Multilingual E5 Text Embeddings: A Technical Report

Liang Wang, Nan Yang, Xiaolong Huang +3 authors

Multilingual E5 text embedding models are released in different sizes and include an instruction-tuned variant, achieving performance comparable to state-of-the-art English-only models.

23contrastive pre-trainingfine-tuningHF ↗arXiv ↗
26

Rethinking Interpretability in the Era of Large Language Models

Chandan Singh, Jeevana Priya Inala, Michel Galley +2 authors

Large language models offer new opportunities in interpretable machine learning, enabling more complex explanations but also introducing challenges like hallucinations and high computational costs.

23large language modelsLLM interpretationHF ↗arXiv ↗
27

Hydragen: High-Throughput LLM Inference with Shared Prefixes

Jordan Juravsky, Bradley Brown, Ryan Ehrlich +3 authors

Hydragen improves LLM inference efficiency by decomposing attention operations into shared prefix and unique suffix computations, reducing memory reads and accelerating matrix multiplications, achieving significant throughput improvements.

21Transformer-based large language modelsLLMsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号