TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

gsm8k 相关论文

20 篇论文 · 按点赞排序

03

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +55 authors

The Qwen2 series, comprising 0.5 to 72 billion parameter models, surpasses prior open models across language understanding, generation, multilingualism, coding, math, and reasoning, with exceptional performance in benchmarks like MMLU, GPQA, HumanEval, GSM8K, BBH, MT-Bench, Arena-Hard, and LiveCodeBench.

175Mixture-of-Expertslanguage modelsHF ↗arXiv ↗
04

Qwen2.5-Omni Technical Report

Jin Xu, Zhifang Guo, Jinzheng He +11 authors

Qwen2.5-Omni is a multimodal model that processes text, images, audio, and video in a streaming fashion and generates text and speech using a dual-track architecture, achieving state-of-the-art performance on multimodal benchmarks.

173block-wise processingTMRoPE (Time-aligned Multimodal RoPE)HF ↗arXiv ↗
05

TAPS: Task Aware Proposal Distributions for Speculative Sampling

Mohamad Zbib, Mohamad Bazzi, Ammar Mohanna +2 authors

Speculative decoding effectiveness depends on draft model training data alignment with downstream tasks, with specialized drafters performing better when combined through confidence-based routing rather than simple averaging.

146speculative decodingdraft modelHF ↗arXiv ↗
07

Large Language Models as Optimizers

Chengrun Yang, Xuezhi Wang, Yifeng Lu +4 authors

OPRO, a method using large language models to optimize tasks described in natural language, outperforms human-designed prompts on various benchmark datasets.

79derivative-based algorithmsOptimization by PROmpting (OPRO)HF ↗arXiv ↗
08

LFM2 Technical Report

Alexander Amini, Anna Banaszak, Harold Benoit +30 authors

LFM2, a family of compact foundation models, achieves high efficiency and performance on-device through hardware-in-the-loop architecture search and advanced training techniques, supporting various tasks including multimodal applications.

65Liquid Foundation Modelshardware-in-the-loop architecture searchHF ↗arXiv ↗
11

Iterative Reasoning Preference Optimization

Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho +3 authors

An iterative preference optimization method using a modified DPO loss improves reasoning accuracy on various datasets by optimizing winning and losing reasoning steps in Chain-of-Thought candidates.

50iterative preference optimizationChain-of-Thought (CoT)HF ↗arXiv ↗
12

MALT: Improving Reasoning with Multi-Agent LLM Training

Sumeet Ramesh Motwani, Chandler Smith, Rocktim Jyoti Das +6 authors

Multi-agent LLM training improves performance on reasoning tasks by assigning specialized roles and utilizing joint outcome-based rewards to enhance collaboration among models.

46sequential multi-agent setupheterogeneous LLMsHF ↗arXiv ↗
19

Learning From Mistakes Makes LLM Better Reasoner

Shengnan An, Zexiong Ma, Zeqi Lin +3 authors

LeMa, a learning-from-mistakes approach, enhances LLMs' mathematical reasoning by learning from inaccurate reasoning paths corrected by GPT-4, surpassing SOTA performance on math problems.

29Large language modelsLearning from MistakesHF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
11 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
13
37 篇论文
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号