TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 28 – May 4, 2025
本周最热99

Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Yiping Wang, Qing Yang, Zhiyuan Zeng +11 authors

Reinforcement learning with verifiable reward using one training example significantly enhances math reasoning capabilities of large language models.

reinforcement learningverifiable reward1-shot RLVRlarge language modelsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Towards Understanding Camera Motions in Any Video

Zhiqiu Lin, Siyuan Cen, Daniel Jiang +12 authors

CameraBench assesses and enhances camera motion understanding by providing a large-scale dataset, taxonomy, and benchmark for evaluating SfM and VLMs, emphasizing the need for both semantic and geometric information.

99Structure-from-Motion (SfM)Video-Language Models (VLMs)HF ↗arXiv ↗
03

The Leaderboard Illusion

Shivalika Singh, Yiyang Nan, Alex Wang +10 authors

Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. We show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on the arena distribution, based on our conservative estimates. Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field

71HF ↗arXiv ↗
04

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Prateek Chhikara, Dev Khant, Saket Aryan +2 authors

Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.

71Mem0memory-centric architectureHF ↗arXiv ↗
08

Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Chris, Yichen Wei, Yi Peng +10 authors

Skywork R1V2 enhances multimodal reasoning through a hybrid reinforcement learning approach that balances reward-model guidance and rule-based strategies, improving training efficiency with the Selective Sample Buffer mechanism and mitigating visual hallucinations.

57hybrid reinforcement learningreward-model guidanceHF ↗arXiv ↗
09

Phi-4-reasoning Technical Report

Marah Abdin, Sahaj Agarwal, Ahmed Awadallah +20 authors

Phi-4-reasoning, a 14-billion parameter model enhanced with supervised fine-tuning and reinforcement learning, outperforms larger models on complex reasoning tasks across various benchmarks.

56Phi-4-reasoningPhi-4-reasoning-plusHF ↗arXiv ↗
10

DeepCritic: Deliberate Critique with Large Language Models

Wenkai Yang, Jingwen Chen, Yankai Lin +1 authors

A novel two-stage framework using Qwen2.5-72B-Instruct enhances LLMs' math critique ability by generating detailed step-wise critiques and applying reinforcement learning, resulting in better error identification and refinement.

54Large Language Models (LLMs)critique modelsHF ↗arXiv ↗
11

ReasonIR: Training Retrievers for Reasoning Tasks

Rulin Shao, Rui Qiao, Varsha Kishore +8 authors

ReasonIR-8B, a retriever trained with synthetic challenging queries, achieves state-of-the-art results in reasoning-intensive information retrieval and RAG tasks, improving performance and computational efficiency.

54retrieversynthetic data generationHF ↗arXiv ↗
14

A Survey of Interactive Generative Video

Jiwen Yu, Yiran Qin, Haoxuan Che +7 authors

Interactive Generative Video (IGV) combines generative capabilities with interactive features to enable diverse applications across gaming, embodied AI, and autonomous driving, focusing on challenges such as real-time generation, control, memory, dynamics, and intelligence.

46Interactive Generative Video (IGV)generative capabilitiesHF ↗arXiv ↗
25

TesserAct: Learning 4D Embodied World Models

Haoyu Zhen, Qiao Sun, Hongxin Zhang +4 authors

The paper introduces a 4D world model trained on RGB-DN videos to predict spatial and temporal dynamics in embodied environments, improving inverse dynamics learning, view synthesis, and policy performance.

224D world modelRGB-DN videosHF ↗arXiv ↗
26

Kimi-Audio Technical Report

KimiTeam, Ding Ding, Zeqian Ju +37 authors

Kimi-Audio, an open-source audio foundation model, achieves state-of-the-art performance across audio-related tasks through a novel LLM-based architecture and comprehensive training and evaluation processes.

21LLM-based architecturecontinuous featuresHF ↗arXiv ↗
29

Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions

Mohammad Mahdi Abootorabi, Omid Ghahroodi, Pardis Sadat Zahraei +17 authors

This survey provides a comprehensive overview of generative AI applications in character animation, covering facial animation, expression rendering, image synthesis, avatar creation, gesture modeling, motion synthesis, object generation, and texture synthesis.

18foundation modelsdiffusion modelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号