TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

September 2025
本月最热665

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

Jeffrey Amico, Gabriel Passamani Andrade, John Donaghy +12 authors

Swarm sAmpling Policy Optimization (SAPO) is a decentralized and asynchronous RL algorithm that enhances post-training language models without supervised fine-tuning, achieving significant reward gains and scalability across diverse hardware.

reinforcement learningpost-traininglanguage modelsDeepSeek-R1-ZeroHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu +22 authors

Agentic reinforcement learning transforms large language models into autonomous decision-making agents by leveraging temporally extended POMDPs, enhancing capabilities like planning and reasoning through reinforcement learning.

238agentic reinforcement learningLLM RLHF ↗arXiv ↗
06

Why Language Models Hallucinate

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala +1 authors

Language models produce incorrect statements due to training and evaluation procedures that reward guessing over acknowledging uncertainty, leading to a need for socio-technical changes in benchmark scoring.

200hallucinationsbinary classificationHF ↗arXiv ↗
08

LongLive: Real-time Interactive Long Video Generation

Shuai Yang, Wei Huang, Ruihang Chu +9 authors

LongLive is a frame-level autoregressive framework for real-time and interactive long video generation, addressing efficiency and quality challenges through causal attention, KV-recache, streaming long tuning, and short window attention.

191frame-level autoregressivediffusion modelsHF ↗arXiv ↗
10

Qwen3-Omni Technical Report

Jin Xu, Zhifang Guo, Hangrui Hu +35 authors

Qwen3-Omni, a multimodal model, achieves state-of-the-art performance across text, image, audio, and video, using a Thinker-Talker MoE architecture and a lightweight causal ConvNet for efficient streaming synthesis.

155multimodal modelThinker-Talker MoE architectureHF ↗arXiv ↗
14

Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed +4 authors

Baseer, a vision-language model fine-tuned for Arabic document OCR, achieves state-of-the-art performance using a decoder-only strategy and a large-scale dataset, outperforming existing solutions with a WER of 0.25.

135Multimodal Large Language Modelsvision-language modelHF ↗arXiv ↗
20

Quantile Advantage Estimation for Entropy-Safe Reasoning

Junkang Wu, Kexin Huang, Jiancan Wu +3 authors

Quantile Advantage Estimation stabilizes reinforcement learning with verifiable rewards by addressing entropy issues and improving performance on large language models.

119Reinforcement Learning with Verifiable Rewardsentropy collapseHF ↗arXiv ↗
22

Scaling Agents via Continual Pre-training

Liangcai Su, Zhen Zhang, Guangyu Li +19 authors

AgentFounder, a deep research agent model incorporating Agentic Continual Pre-training, achieves state-of-the-art performance in agentic tasks while maintaining strong tool-use ability.

118Large language modelsagentic systemsHF ↗arXiv ↗
27

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Tong Zheng, Hongming Zhang, Wenhao Yu +8 authors

Parallel-R1, a reinforcement learning framework, enhances large language models' reasoning capabilities by enabling parallel thinking through a progressive curriculum, leading to significant performance improvements on math benchmarks.

106parallel thinkingreinforcement learningHF ↗arXiv ↗
29

LIMI: Less is More for Agency

Yang Xiao, Mohan Jiang, Jie Sun +18 authors

LIMI demonstrates that sophisticated agentic intelligence can emerge from minimal, strategically curated demonstrations, outperforming data-intensive models on agency benchmarks.

104Agencyautonomous agentsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号