TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 21 – Jul 27, 2025
本周最热323

Group Sequence Policy Optimization

Chujie Zheng, Shixuan Liu, Mingze Li +9 authors

Group Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.

Group Sequence Policy OptimizationGSPOreinforcement learninglarge language modelsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

GUI-G^2: Gaussian Reward Modeling for GUI Grounding

Fei Tang, Zhangxuan Gu, Zhengxi Lu +9 authors

A new reward framework, GUI-G$^2$, models GUI elements as continuous Gaussian distributions to improve autonomous interaction through dense gradient signals, outperforming existing methods in spatial reasoning tasks.

135reinforcement learningbinary rewardsHF ↗arXiv ↗
05

nablaNABLA: Neighborhood Adaptive Block-Level Attention

Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva +6 authors

NABLA, a Neighborhood Adaptive Block-Level Attention mechanism, enhances video diffusion transformers by reducing computational overhead without significantly impacting generative quality or visual fidelity.

103transformer-based architecturesvideo generationHF ↗arXiv ↗
06

Yume: An Interactive World Generation Model

Xiaofeng Mao, Shaoheng Lin, Zhen Li +7 authors

A framework for generating and exploring interactive, high-fidelity video worlds from images using a Masked Video Diffusion Transformer, advanced sampling techniques, and model acceleration.

92camera motion quantizationMasked Video Diffusion TransformerHF ↗arXiv ↗
07

The Invisible Leash: Why RLVR May Not Escape Its Origin

Fang Wu, Weihao Xuan, Ximing Lu +2 authors

Reinforcement Learning with Verifiable Rewards (RLVR) enhances precision but may limit exploration and discovery of new solutions, suggesting potential limits to its effectiveness in expanding reasoning capabilities.

85Reinforcement Learning with Verifiable RewardsRLVRHF ↗arXiv ↗
08

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu +106 authors

Step-Audio~2, an end-to-end multi-modal large language model, integrates latent audio encoding and reinforcement learning to achieve state-of-the-art performance in ASR, audio understanding, and speech conversation, incorporating discrete audio token generation and retrieval-augmented generation.

77latent audio encoderreasoning-centric reinforcement learningHF ↗arXiv ↗
16

GR-3 Technical Report

Chilam Cheang, Sijin Chen, Zhongren Cui +18 authors

GR-3, a large-scale vision-language-action model, demonstrates exceptional generalization and fine-tuning capabilities, excelling in complex tasks and outperforming state-of-the-art methods.

47vision-language-action modelco-trainingHF ↗arXiv ↗
19

Captain Cinema: Towards Short Movie Generation

Junfei Xiao, Ceyuan Yang, Lvmin Zhang +7 authors

Captain Cinema generates coherent short movies from textual descriptions using top-down keyframe planning and bottom-up video synthesis with Multimodal Diffusion Transformers.

42top-down keyframe planningbottom-up video synthesisHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号