TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本年最热665

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

Jeffrey Amico, Gabriel Passamani Andrade, John Donaghy +12 authors

Swarm sAmpling Policy Optimization (SAPO) is a decentralized and asynchronous RL algorithm that enhances post-training language models without supervised fine-tuning, achieving significant reward gains and scalability across diverse hardware.

reinforcement learningpost-traininglanguage modelsDeepSeek-R1-ZeroHF ↗arXiv ↗

606 篇论文 · 按点赞排序

06

Qwen3 Technical Report

An Yang, Anfeng Li, Baosong Yang +57 authors

Qwen3, a unified series of large language models, integrates thinking and non-thinking modes, reduces computational resources, and achieves state-of-the-art performance across various tasks and languages.

343large language modelsdense architectureHF ↗arXiv ↗
07

mHC: Manifold-Constrained Hyper-Connections

Zhenda Xie, Yixuan Wei, Huanqi Cao +16 authors

Manifold-Constrained Hyper-Connections (mHC) stabilize and scale residual connection architectures by restoring identity mapping properties through manifold projection and infrastructure optimization.

330Hyper-Connections (HC)Manifold-Constrained Hyper-Connections (mHC)HF ↗arXiv ↗
08

Group Sequence Policy Optimization

Chujie Zheng, Shixuan Liu, Mingze Li +9 authors

Group Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.

320Group Sequence Policy OptimizationGSPOHF ↗arXiv ↗
09

DINOv3

Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23 authors

DINOv3, a self-supervised learning model, achieves superior performance across various vision tasks by scaling datasets and models, addressing dense feature degradation, and enhancing flexibility with post-hoc strategies.

311self-supervised learningDINOv3HF ↗arXiv ↗
13

MiniMax-01: Scaling Foundation Models with Lightning Attention

MiniMax, Aonian Li, Bangwei Gong +87 authors

The MiniMax-01 series, including MiniMax-Text-01 and MiniMax-VL-01, offer superior long-context processing and match state-of-the-art model performance with longer context windows through lightning attention, Mixture of Experts (MoE), and efficient parallel strategies.

304lightning attentionMixture of ExpertsHF ↗arXiv ↗
17

Agent Learning via Early Experience

Kai Zhang, Xiangchao Chen, Bo Liu +27 authors

Early experience, using agent-generated interaction data without reward signals, improves policy effectiveness and generalization, serving as a bridge between imitation learning and reinforcement learning.

276reinforcement learningearly experienceHF ↗arXiv ↗
18

Qwen-Image Technical Report

Chenfei Wu, Jiahao Li, Jingren Zhou +36 authors

Qwen-Image, an image generation model, advances text rendering and image editing through a comprehensive data pipeline, progressive training, and dual-encoding mechanism.

276data pipelineprogressive trainingHF ↗arXiv ↗
19

Intern-S1: A Scientific Multimodal Foundation Model

Lei Bai, Zhongrui Cai, Maosong Cao +172 authors

Intern-S1, a multimodal Mixture-of-Experts model with extensive pre-training and reinforcement learning, achieves top-tier performance in general reasoning and outperforms closed-source models in scientific tasks.

274Mixture-of-Experts (MoE)reinforcement learning (RL)HF ↗arXiv ↗
21

Reinforcement Pre-Training

Qingxiu Dong, Li Dong, Yao Tang +4 authors

Reinforcement Pre-Training (RPT) improves language model accuracy through reinforcement learning and offers a scalable method for leveraging text data for general-purpose RL.

265Reinforcement Pre-Training (RPT)next-token predictionHF ↗arXiv ↗
29

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu +22 authors

Agentic reinforcement learning transforms large language models into autonomous decision-making agents by leveraging temporally extended POMDPs, enhancing capabilities like planning and reasoning through reinforcement learning.

239agentic reinforcement learningLLM RLHF ↗arXiv ↗
30

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22 authors

Kandinsky 5.0 is a family of state-of-the-art generative models for high-resolution images and short videos, featuring model lineups with varying parameters and enhanced training techniques to achieve superior quality and performance.

236foundation modelshigh-resolution image synthesisHF ↗arXiv ↗
1 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号