TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热463

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI, Daya Guo, Dejian Yang +197 authors

DeepSeek-R1-Zero and DeepSeek-R1 utilize reinforcement learning and multi-stage training to enhance reasoning capabilities, with DeepSeek-R1 achieving performance comparable to OpenAI-o1-1217.

reinforcement learningmulti-stage trainingcold-start dataQwenHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

MiniMax-01: Scaling Foundation Models with Lightning Attention

MiniMax, Aonian Li, Bangwei Gong +87 authors

The MiniMax-01 series, including MiniMax-Text-01 and MiniMax-VL-01, offer superior long-context processing and match state-of-the-art model performance with longer context windows through lightning attention, Mixture of Experts (MoE), and efficient parallel strategies.

304lightning attentionMixture of ExpertsHF ↗arXiv ↗
04

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Kimi Team, Angang Du, Bofei Gao +91 authors

A multi-modal LLM trained with reinforcement learning achieves state-of-the-art reasoning performance across various benchmarks by utilizing long context scaling and effective policy optimization methods.

132next token predictionreinforcement learningHF ↗arXiv ↗
06

Evolving Deeper LLM Thinking

Kuang-Huei Lee, Ian Fischer, Yueh-Hua Wu +4 authors

Mind Evolution, an evolutionary search strategy using a language model, outperforms other inference methods in natural language planning tasks by generating, recombining, and refining candidate responses.

116evolutionary search strategyLarge Language ModelsHF ↗arXiv ↗
09

Search-o1: Agentic Search-Enhanced Large Reasoning Models

Xiaoxi Li, Guanting Dong, Jiajie Jin +5 authors

Search-o1 enhances large reasoning models with an agentic retrieval-augmented generation mechanism and a Reason-in-Documents module to improve performance on complex reasoning tasks.

106Large reasoning modelsreinforcement learningHF ↗arXiv ↗
16

Tensor Product Attention Is All You Need

Yifan Zhang, Yifeng Liu, Huizhuo Yuan +4 authors

T6, a new Transformer architecture using Tensor Product Attention, improves language model performance and memory efficiency, allowing for longer sequence processing.

91Tensor Product AttentionTPAHF ↗arXiv ↗
20

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Yilun Zhao, Lujing Xie, Haowei Zhang +16 authors

A comprehensive benchmark for evaluating foundation models in video understanding includes domain-specific questions and expert evaluations, highlighting the gap between current models and human expertise.

82multimodal foundation modelsexpert-level reasoningHF ↗arXiv ↗
21

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1155 authors

HLE is a challenging multi-modal benchmark that highlights the limitations of current LLMs in closed-ended academic questions.

78large language model (LLM)benchmarksHF ↗arXiv ↗
24

Qwen2.5-1M Technical Report

An Yang, Bowen Yu, Chengyuan Li +25 authors

The Qwen2.5-1M series models extend context length to 1 million tokens with enhanced long-context capabilities, employing techniques like long data synthesis and progressive pre-training, and are supported by an open-source inference framework with sparse attention and kernel optimizations.

72long-context pre-traininglong data synthesisHF ↗arXiv ↗
26

Enabling Scalable Oversight via Self-Evolving Critic

Zhengyang Tang, Ziniu Li, Zhenyang Xiao +8 authors

SCRIT, a self-evolving critique framework, enhances LLMs' critique capabilities using synthetic data and self-validation, achieving significant improvements in critique-correction and error identification benchmarks.

72Large Language Models (LLMs)self-evolvingHF ↗arXiv ↗
27

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev

Shared Recurrent Memory Transformer (SRMT) enhances cooperation in multi-agent reinforcement learning through implicit information exchange, outperforming baselines in navigation tasks and generalizing to unseen scenarios.

70multi-agent reinforcement learningMARLHF ↗arXiv ↗
29

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Omkar Thawakar, Dinura Dissanayake, Ketan More +12 authors

A framework for evaluating and improving step-by-step visual reasoning in large language models using a specialized benchmark and a novel multimodal model trained with curriculum learning.

67visual reasoninglarge language modelsHF ↗arXiv ↗
30

UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Yujia Qin, Yining Ye, Junjie Fang +32 authors

UI-TARS, a native GUI agent model using screenshots as input, outperforms commercial models in various benchmarks through enhanced perception, unified action modeling, system-2 reasoning, and iterative training with reflective online traces.

64native GUI agent modelcontext-aware understandingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号