TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 20 – Jan 26, 2025
本周最热463

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI, Daya Guo, Dejian Yang +197 authors

DeepSeek-R1-Zero and DeepSeek-R1 utilize reinforcement learning and multi-stage training to enhance reasoning capabilities, with DeepSeek-R1 achieving performance comparable to OpenAI-o1-1217.

reinforcement learningmulti-stage trainingcold-start dataQwenHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Kimi Team, Angang Du, Bofei Gao +91 authors

A multi-modal LLM trained with reinforcement learning achieves state-of-the-art reasoning performance across various benchmarks by utilizing long context scaling and effective policy optimization methods.

132next token predictionreinforcement learningHF ↗arXiv ↗
03

Evolving Deeper LLM Thinking

Kuang-Huei Lee, Ian Fischer, Yueh-Hua Wu +4 authors

Mind Evolution, an evolutionary search strategy using a language model, outperforms other inference methods in natural language planning tasks by generating, recombining, and refining candidate responses.

116evolutionary search strategyLarge Language ModelsHF ↗arXiv ↗
06

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Yilun Zhao, Lujing Xie, Haowei Zhang +16 authors

A comprehensive benchmark for evaluating foundation models in video understanding includes domain-specific questions and expert evaluations, highlighting the gap between current models and human expertise.

82multimodal foundation modelsexpert-level reasoningHF ↗arXiv ↗
08

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev

Shared Recurrent Memory Transformer (SRMT) enhances cooperation in multi-agent reinforcement learning through implicit information exchange, outperforming baselines in navigation tasks and generalizing to unseen scenarios.

70multi-agent reinforcement learningMARLHF ↗arXiv ↗
10

UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Yujia Qin, Yining Ye, Junjie Fang +32 authors

UI-TARS, a native GUI agent model using screenshots as input, outperforms commercial models in various benchmarks through enhanced perception, unified action modeling, system-2 reasoning, and iterative training with reflective online traces.

64native GUI agent modelcontext-aware understandingHF ↗arXiv ↗
17

Autonomy-of-Experts Models

Ang Lv, Ruobing Xie, Yining Qian +5 authors

Autonomy-of-Experts (AoE) improves MoE models by enabling experts to autonomously select inputs based on self-evaluated capacity, reducing the need for routers and enhancing performance.

44Mixture-of-ExpertsMoEHF ↗arXiv ↗
20

Improving Video Generation with Human Feedback

Jie Liu, Gongye Liu, Jiajun Liang +15 authors

VideoReward and alignment algorithms improve video generation by addressing unsmooth motion and misalignment using human feedback in a unified reinforcement learning framework.

33rectified flow techniqueshuman preference datasetHF ↗arXiv ↗
21

Reasoning Language Models: A Blueprint

Maciej Besta, Julia Barth, Eric Schreiber +15 authors

A comprehensive blueprint for modularizing reasoning language models (RLMs) to enhance accessibility and scalability, incorporating diverse reasoning structures, RL concepts, and supervision schemes.

33reasoning language modelsRLMsHF ↗arXiv ↗
25

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Zhongwei Ren, Yunchao Wei, Xun Guo +4 authors

VideoWorld, an auto-regressive video generation model trained from visual data, acquires complex knowledge including rules, reasoning, and planning, and achieves high performance in tasks like Video-Go and robotic control without traditional reinforcement learning techniques.

27auto-regressive video generationVideoWorldHF ↗arXiv ↗
30

Video Depth Anything: Consistent Depth Estimation for Super-Long Videos

Sili Chen, Hengkai Guo, Shengnan Zhu +4 authors

Video Depth Anything provides high-quality, consistent depth estimation for long videos using an efficient spatial-temporal head and a temporal consistency loss, achieving state-of-the-art performance in zero-shot video depth estimation.

23monocular depth estimationtemporal inconsistencyHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号