TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

gpt-4 相关论文

49 篇论文 · 按点赞排序

21

Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar +3 authors

Orca, a 13-billion parameter model, enhances small models by imitating the reasoning process from large foundation models using rich signals and diverse data, surpassing existing models in complex reasoning benchmarks.

51imitation learninglarge foundation modelsHF ↗arXiv ↗
31

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

David Rein, Betty Li Hou, Asa Cooper Stickland +5 authors

A dataset of extremely difficult multiple-choice questions challenges both experts and AI systems, facilitating the development of scalable oversight methods for AI-generated knowledge.

37GPQAmultiple-choice questionsHF ↗arXiv ↗
32

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97 authors

InternLM2 is an open-source LLM that outperforms predecessors through innovative pre-training and optimization techniques, including Supervised Fine-Tuning and Conditional Online Reinforcement Learning from Human Feedback.

35Large Language ModelsLLMsHF ↗arXiv ↗
33

JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Lianghui Zhu, Xinggang Wang, Xinlong Wang

Large Language Models fine-tuned as scalable judges (JudgeLM) achieve state-of-the-art performance in evaluating open-ended benchmarks through a comprehensive dataset and benchmark, enhancing judgment efficiency and accuracy.

35Large Language Modelsfine-tuningHF ↗arXiv ↗
35

Shepherd: A Critic for Language Model Generation

Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan +7 authors

Shepherd, a small language model tuned for critique, outperforms or ties with larger models like ChatGPT in refining language model outputs using a high-quality feedback dataset.

33language modelcritiqueHF ↗arXiv ↗
37

Aligning Large Multimodal Models with Factually Augmented RLHF

Zhiqing Sun, Sheng Shen, Shengcao Cao +9 authors

The paper presents Factually Augmented RLHF to address multimodal misalignment in large multimodal models, significantly improving hallucination reduction and overall performance on vision-language tasks compared to existing methods.

32Reinforcement Learning from Human Feedback (RLHF)vision-language alignmentHF ↗arXiv ↗
40

Learning From Mistakes Makes LLM Better Reasoner

Shengnan An, Zexiong Ma, Zeqi Lin +3 authors

LeMa, a learning-from-mistakes approach, enhances LLMs' mathematical reasoning by learning from inaccurate reasoning paths corrected by GPT-4, surpassing SOTA performance on math problems.

29Large language modelsLearning from MistakesHF ↗arXiv ↗
2 / 3

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号