TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 13 – Apr 19, 2026

50 篇论文 · 按点赞排序

36

SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks

Tianyi Wang, Yixia Li, Long Li +6 authors

Sequence-Level PPO addresses instability in long-chain-of-thought reasoning by reformulating the process as a contextual bandit problem with decoupled value functions for improved efficiency.

29Proximal Policy OptimizationLarge Language ModelsHF ↗arXiv ↗
38

Multi-User Large Language Model Agents

Shu Yang, Shenzhe Zhu, Hao Zhu +5 authors

Multi-user large language model agents face challenges in handling conflicting objectives, privacy preservation, and coordination efficiency in multi-principal decision-making scenarios.

29large language modelsmulti-user interactionHF ↗arXiv ↗
40

SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Binbin Zheng, Xing Ma, Yiheng Liang +6 authors

SCOPE enhances on-policy distillation by adapting supervision paths based on trajectory correctness, using teacher-perplexity-weighted KL distillation for incorrect trajectories and student-perplexity-weighted MLE for correct ones, achieving superior reasoning performance.

27on-policy reinforcement learningtoken-level credit assignmentHF ↗arXiv ↗
43

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

Jiacheng Liu, Xiaohan Zhao, Xinyi Shang +1 authors

The study analyzes Claude Code's architecture, identifying five motivating human values and tracing them through thirteen design principles to specific implementation choices, including a core while-loop architecture and supporting systems for safety, context management, and extensibility.

25agentic coding toolshell commandsHF ↗arXiv ↗
45

Introspective Diffusion Language Models

Yifan Yu, Yuqing Jian, Junxiong Wang +12 authors

Introspective Diffusion Language Models address quality gaps with autoregressive models by enforcing introspective consistency through novel decoding algorithms and optimized inference engines.

25diffusion language modelsautoregressive modelsHF ↗arXiv ↗
46

ELT: Elastic Looped Transformers for Visual Generation

Sahil Goyal, Swayam Agrawal, Gautham Govind Anil +3 authors

Elastic Looped Transformers utilize recurrent transformer architecture with weight-sharing and intra-loop self-distillation to achieve parameter-efficient visual generation with adjustable computational cost and generation quality.

24recurrent transformer architectureweight-sharingHF ↗arXiv ↗
49

Target Policy Optimization

Jean Kaddour

Target Policy Optimization separates policy update decisions from probability assignment in reinforcement learning, improving performance over standard policy gradient methods in sparse reward scenarios.

23policy-gradient methodspolicy optimizationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号