TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

31

The Leaderboard Illusion

Shivalika Singh, Yiyang Nan, Alex Wang +10 authors

Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. We show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on the arena distribution, based on our conservative estimates. Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field

71HF ↗arXiv ↗
32

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Prateek Chhikara, Dev Khant, Saket Aryan +2 authors

Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.

71Mem0memory-centric architectureHF ↗arXiv ↗
33

Seedream 3.0 Technical Report

Yu Gao, Lixue Gong, Qiushan Guo +28 authors

Seedream 3.0 improves Chinese-English bilingual image generation by enhancing data training, pre-training techniques, and post-training aesthetics, resulting in higher visual quality and faster image generation.

71defect-aware trainingdual-axis collaborative data-samplingHF ↗arXiv ↗
38

Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Zhenyi Liao, Qingsong Xie, Yanhao Zhang +4 authors

The study enhances visual-spatial reasoning in multi-modal large language models through GRPO training using the VSI-100k dataset, demonstrating significant performance improvements over base models.

67multi-modal large language modelsvisual-spatial intelligenceHF ↗arXiv ↗
39

Describe Anything: Detailed Localized Image and Video Captioning

Long Lian, Yifan Ding, Yunhao Ge +8 authors

The Describe Anything Model (DAM) leverages a focal prompt and localized vision backbone to achieve detailed localized captioning, outperforming existing models on various benchmarks through a semi-supervised data pipeline.

66focal promptlocalized vision backboneHF ↗arXiv ↗
41

An Empirical Study of GPT-4o Image Generation Capabilities

Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16 authors

An empirical study of GPT-4o's image generation capabilities across multiple tasks reveals its strengths and limitations compared to other models, highlighting the importance of architectural design and data scaling in unified generative frameworks.

64GANdiffusion modelsHF ↗arXiv ↗
42

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Jiazhan Feng, Shijue Huang, Xingwei Qu +6 authors

ReTool, a tool-integrated learning framework, enhances reasoning models with real-time code execution and reinforcement learning, significantly improving performance in structured problem-solving tasks like mathematical reasoning.

63reasoning modelsreinforcement learningHF ↗arXiv ↗
48

Antidistillation Sampling

Yash Savani, Asher Trockman, Zhili Feng +4 authors

Antidistillation sampling modifies a model's next-token probability distribution to disrupt the generation of reasoning traces for distillation without affecting model performance.

60antidistillation samplingnext-token probability distributionHF ↗arXiv ↗
49

Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Chris, Yichen Wei, Yi Peng +10 authors

Skywork R1V2 enhances multimodal reasoning through a hybrid reinforcement learning approach that balances reward-model guidance and rule-based strategies, improving training efficiency with the Selective Sample Buffer mechanism and mitigating visual hallucinations.

59hybrid reinforcement learningreward-model guidanceHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号