TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 21 – Apr 27, 2025
本周最热141

Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Yang Yue, Zhiqi Chen, Rui Lu +5 authors

Despite initial claims, reinforcement learning with verifiable rewards does not introduce fundamentally new reasoning abilities to LLMs, instead it enhances performance by biasing output distribution toward rewarded paths without expanding the reasoning boundary.

Reinforcement Learning with Verifiable RewardsRLVRpass\@kreasoning pathsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

TTRL: Test-Time Reinforcement Learning

Yuxin Zuo, Kaiyan Zhang, Shang Qu +7 authors

Test-Time Reinforcement Learning (TTRL) enhances Large Language Models (LLMs) using unlabeled data through reinforcement learning, improving performance across tasks.

123Reinforcement Learning (RL)Large Language Models (LLMs)HF ↗arXiv ↗
06

Learning to Reason under Off-Policy Guidance

Jianhao Yan, Yafu Li, Zican Hu +5 authors

LUFFY enhances zero-RL models with off-policy guidance, improving reasoning and generalization through balanced imitation and exploration.

88large reasoning modelsreinforcement learningHF ↗arXiv ↗
09

Describe Anything: Detailed Localized Image and Video Captioning

Long Lian, Yifan Ding, Yunhao Ge +8 authors

The Describe Anything Model (DAM) leverages a focal prompt and localized vision backbone to achieve detailed localized captioning, outperforming existing models on various benchmarks through a semi-supervised data pipeline.

66focal promptlocalized vision backboneHF ↗arXiv ↗
11

Tina: Tiny Reasoning Models via LoRA

Shangshang Wang, Julian Asilis, Ömer Faruk Akgül +3 authors

Tina, a family of tiny reasoning models, achieves high reasoning performance at minimal computational cost through LoRA during reinforcement learning.

56reinforcement learningLoRAHF ↗arXiv ↗
14

ToolRL: Reward is All Tool Learning Needs

Cheng Qian, Emre Can Acikgoz, Qi He +5 authors

A comprehensive study on reward design for tool use within reinforcement learning improves LLMs' tool use capabilities and generalization performance over supervised fine-tuning.

49supervised fine-tuningreinforcement learningHF ↗arXiv ↗
15

FlowReasoner: Reinforcing Query-Level Meta-Agents

Hongcheng Gao, Yue Liu, Yufei He +6 authors

A meta-agent named FlowReasoner automates the design of query-level multi-agent systems using DeepSeek R1 and reinforcement learning, excelling in performance, complexity, and efficiency across benchmarks.

47DeepSeek R1reinforcement learningHF ↗arXiv ↗
17

NodeRAG: Structuring Graph-based RAG with Heterogeneous Nodes

Tianyang Xu, Haojie Zheng, Chengze Li +4 authors

NodeRAG, a graph-centric framework with heterogeneous graph structures, enhances Retrieval-augmented Generation (RAG) by improving integration and performance in indexing, querying, and question-answering.

44Retrieval-augmented generationNodeRAGHF ↗arXiv ↗
23

Trillion 7B Technical Report

Sungjun Han, Juyoung Suk, Suyeong An +5 authors

Trillion-7B is a highly efficient multilingual LLM leveraging Cross-lingual Document Attention (XLDA) for knowledge transfer and achieving competitive performance with minimal multilingual training data.

32Cross-lingual Document Attention (XLDA)HF ↗arXiv ↗
24

I-Con: A Unifying Framework for Representation Learning

Shaden Alshammari, John Hershey, Axel Feldmann +2 authors

A unified framework using KL divergence between supervisory and learned representations generalizes multiple machine learning loss functions and improves unsupervised image classification and debiasing.

31information-theoretic equationKL divergenceHF ↗arXiv ↗
27

UFO2: The Desktop AgentOS

Chaoyun Zhang, He Huang, Chiming Ni +18 authors

UFO2, an AgentOS for Windows, integrates multimodal LLMs with native APIs and hybrid control detection to enhance robust and scalable desktop automation.

29multimodal large language modelsAgentOSHF ↗arXiv ↗
30

AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Ivan Moshkov, Darragh Hanley, Ivan Sorokin +5 authors

This paper presents our winning submission to the AI Mathematical Olympiad - Progress Prize 2 (AIMO-2) competition. Our recipe for building state-of-the-art mathematical reasoning models relies on three key pillars. First, we create a large-scale dataset comprising 540K unique high-quality math problems, including olympiad-level problems, and their 3.2M long-reasoning solutions. Second, we develop a novel method to integrate code execution with long reasoning models through iterative training, generation, and quality filtering, resulting in 1.7M high-quality Tool-Integrated Reasoning solutions. Third, we create a pipeline to train models to select the most promising solution from many candidates. We show that such generative solution selection (GenSelect) can significantly improve upon majority voting baseline. Combining these ideas, we train a series of models that achieve state-of-the-art results on mathematical reasoning benchmarks. To facilitate further research, we release our code, models, and the complete OpenMathReasoning dataset under a commercially permissive license.

26HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号