TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

grpo 相关论文

39 篇论文 · 按点赞排序

01

Group Sequence Policy Optimization

Chujie Zheng, Shixuan Liu, Mingze Li +9 authors

Group Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.

320Group Sequence Policy OptimizationGSPOHF ↗arXiv ↗
02

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

Mind Lab, Song Cao, Vic Cao +59 authors

MinT is a managed infrastructure system that enables efficient low-rank adaptation training and serving by keeping base models resident and moving lightweight adapter revisions, scaling across multiple dimensions including large model architectures, reduced storage requirements, and distributed policy management.

223Low-Rank AdaptationLoRAHF ↗arXiv ↗
03

On-Policy Self-Distillation without Any Supervision

Yijiang Li, Bingyang Wang, Yijun Liang +3 authors

Unsupervised on-policy self-distillation improves large language models by using internal consistency and majority-vote pseudo-solutions to correct confident errors without external supervision.

219on-policy self-distillationself-consistencyHF ↗arXiv ↗
05

Your Group-Relative Advantage Is Biased

Fengkai Yang, Zherui Chen, Xiaohan Wang +10 authors

Group-based reinforcement learning from verifier rewards suffers from biased advantage estimation that underestimates hard prompts and overestimates easy prompts, which is addressed through a history-aware adaptive difficulty weighting method that improves performance on mathematical reasoning benchmarks.

158Reinforcement Learning from Verifier Rewardsgroup-based methodsHF ↗arXiv ↗
08

Quantile Advantage Estimation for Entropy-Safe Reasoning

Junkang Wu, Kexin Huang, Jiancan Wu +3 authors

Quantile Advantage Estimation stabilizes reinforcement learning with verifiable rewards by addressing entropy issues and improving performance on large language models.

119Reinforcement Learning with Verifiable Rewardsentropy collapseHF ↗arXiv ↗
09

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia +39 authors

Ovis2.5, a native-resolution vision transformer with multimodal reasoning, achieves state-of-the-art performance on various benchmarks through advanced training techniques and efficient scaling methods.

116vision transformernative-resolutionHF ↗arXiv ↗
10

Self-Distilled Agentic Reinforcement Learning

Zhengxi Lu, Zhiyuan Yao, Zhuowen Han +8 authors

SDAR enhances reinforcement learning for multi-turn agent training by integrating self-distillation through a sigmoid gate that selectively strengthens positive token-level guidance while mitigating negative teacher rejections.

116Reinforcement learningon-policy self-distillationHF ↗arXiv ↗
12

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

Shuang Chen, Kaituo Feng, Hangting Chen +7 authors

OpenSearch-VL presents an open-source framework for training advanced multimodal search agents using reinforcement learning, featuring specialized data curation, diverse tool environments, and a novel training algorithm that improves performance across multiple benchmarks.

106multimodal search agentsagentic reinforcement learningHF ↗arXiv ↗
14

Flow-OPD: On-Policy Distillation for Flow Matching Models

Zhen Fang, Wenxuan Huang, Yu Zeng +8 authors

Flow-OPD addresses limitations in Flow Matching text-to-image models through a two-stage alignment approach combining on-policy distillation and manifold anchor regularization, achieving significant improvements in generation quality and alignment metrics.

102Flow Matchingon-policy distillationHF ↗arXiv ↗
17

Table-R1: Inference-Time Scaling for Table Reasoning

Zheyuan Yang, Lyuhao Chen, Arman Cohan +1 authors

Two post-training strategies, distillation and RLVR, enable inference-time scaling in table reasoning tasks, resulting in a model (Table-R1-Zero) that matches GPT-4.1's performance using fewer parameters and shows strong generalization.

93distillationreinforcement learningHF ↗arXiv ↗
19

GEM: A Gym for Agentic LLMs

Zichen Liu, Anya Sims, Keyu Duan +16 authors

GEM, an open-source environment simulator, facilitates experience-based learning for large language models by providing a standardized framework and diverse environments for training and benchmarking reinforcement learning algorithms.

92large language modelsexperience-based learningHF ↗arXiv ↗
1 / 2

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
12 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
38 篇论文
13
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号