TensorX

Trends · 研究趋势

数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。

返回趋势

on-policy distillation 相关论文

17 篇论文 · 按点赞排序

02

Self-Distilled RLVR

Chenxu Yang, Chuanyu Qin, Qingyi Si +7 authors

RLSD combines reinforcement learning with verifiable rewards and self-distillation to achieve stable training with fine-grained updates and reliable policy direction from environmental feedback.

179on-policy distillationon-policy self-distillationHF ↗arXiv ↗
04

DOPD: Dual On-policy Distillation

Xinlei Yu, Gen Li, Qingyi Si +13 authors

DOPD addresses privilege illusion in on-policy distillation by dynamically routing token-level supervision between teacher and student policies based on advantage gaps and probabilities, improving capability transfer in large and vision-language models.

114on-policy distillationtoken-level signalsHF ↗arXiv ↗
05

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Yaxuan Li, Yuxin Zuo, Bingxiang He +8 authors

On-policy distillation dynamics in large language models depend on compatible thinking patterns between teacher and student models, with successful distillation characterized by alignment on high-probability tokens and requiring teachers to provide novel capabilities beyond student training data.

113on-policy distillationlarge language modelsHF ↗arXiv ↗
08

Flow-OPD: On-Policy Distillation for Flow Matching Models

Zhen Fang, Wenxuan Huang, Yu Zeng +8 authors

Flow-OPD addresses limitations in Flow Matching text-to-image models through a two-stage alignment approach combining on-policy distillation and manifold anchor regularization, achieving significant improvements in generation quality and alignment metrics.

102Flow Matchingon-policy distillationHF ↗arXiv ↗
11

On the Geometry of On-Policy Distillation

Zhennan Shen, Yanshu Li, Qingyu Yin +6 authors

On-policy distillation exhibits distinct parameter space dynamics characterized by relaxed off-principal updates and subspace locking, forming a unique geometric pattern separate from supervised fine-tuning and reinforcement learning with verifiable rewards.

75on-policy distillationparameter spaceHF ↗arXiv ↗
12

Intern-S2-Preview: Scientific Agentic Foundation Model

Lei Bai, Jiaqi Cao, Chiyu Chen +122 authors

Intern-S2-Preview is a scientific agentic foundation model series that integrates multimodal pre-training, multi-task reinforcement learning, and memory-augmented extensions to support long-horizon scientific reasoning and forecasting.

70multimodal pre-trainingsupervised fine-tuningHF ↗arXiv ↗
15

Miles v0.1: Production-Level Post-Training

RadixArk, Tom Chen, Mao Cheng +10 authors

Miles is an open-source, production-ready system for large-scale reinforcement learning and post-training that supports diverse backends, weight synchronization, LoRA, distillation, and diffusion models.

52reinforcement-learningrollout enginesHF ↗arXiv ↗
17

Cliff: Learning Process Rewards from the First Mistake

Peixuan Han, Runhui Wang, Ketan Ramaneti +3 authors

Cliff improves reinforcement learning with verifiable rewards by using an off-the-shelf language model to detect the first reasoning error and shaping token-level advantages accordingly.

15reinforcement learning with verifiable rewardsprocess reward modelingHF ↗arXiv ↗

上升最快

近 6 个月
1
35 篇论文
2
llmNEW
34 篇论文
3
29 篇论文
4
26 篇论文
5
ditNEW
12 篇论文
6
12 篇论文
7
12 篇论文
8
11 篇论文
9
11 篇论文
10
11 篇论文
11
11 篇论文
12
10 篇论文
13
10 篇论文
14
10 篇论文
15
10 篇论文
16
26 篇论文
17
74 篇论文
19
rlvr+200%
13 篇论文
20
12 篇论文

最热方向

按总量
1
3
167 篇论文
5
75 篇论文
6
74 篇论文
10
49 篇论文
11
39 篇论文
12
13
37 篇论文
14
15
35 篇论文
16
34 篇论文
17
33 篇论文
18
29 篇论文
19
29 篇论文
20
29 篇论文
21
28 篇论文
23
27 篇论文
24
27 篇论文
25
26 篇论文
26
29
25 篇论文
30
24 篇论文
31
23 篇论文
32
23 篇论文
33
23 篇论文
34
22 篇论文
35
22 篇论文
36
20 篇论文
37
20 篇论文
38
20 篇论文
39
20 篇论文
40
20 篇论文
41
19 篇论文
42
19 篇论文
43
19 篇论文
46
18 篇论文
48
51
53
55
16 篇论文
56
16 篇论文
58
59
16 篇论文
60

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号