TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

453

Seedream 3.0 Technical Report

Yu Gao, Lixue Gong, Qiushan Guo +28 authors

Seedream 3.0 improves Chinese-English bilingual image generation by enhancing data training, pre-training techniques, and post-training aesthetics, resulting in higher visual quality and faster image generation.

71defect-aware trainingdual-axis collaborative data-samplingHF ↗arXiv ↗
454

The Leaderboard Illusion

Shivalika Singh, Yiyang Nan, Alex Wang +10 authors

Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. We show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on the arena distribution, based on our conservative estimates. Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field

71HF ↗arXiv ↗
455

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev

Shared Recurrent Memory Transformer (SRMT) enhances cooperation in multi-agent reinforcement learning through implicit information exchange, outperforming baselines in navigation tasks and generalizing to unseen scenarios.

70multi-agent reinforcement learningMARLHF ↗arXiv ↗
456

Back to Basics: Let Denoising Generative Models Denoise

Tianhong Li, Kaiming He

Today's denoising diffusion models do not "denoise" in the classical sense, i.e., they do not directly predict clean images. Rather, the neural networks predict noise or a noised quantity. In this paper, we suggest that predicting clean data and predicting noised quantities are fundamentally different. According to the manifold assumption, natural data should lie on a low-dimensional manifold, whereas noised quantities do not. With this assumption, we advocate for models that directly predict clean data, which allows apparently under-capacity networks to operate effectively in very high-dimensional spaces. We show that simple, large-patch Transformers on pixels can be strong generative models: using no tokenizer, no pre-training, and no extra loss. Our approach is conceptually nothing more than "Just image Transformers", or JiT, as we call it. We report competitive results using JiT with large patch sizes of 16 and 32 on ImageNet at resolutions of 256 and 512, where predicting high-dimensional noised quantities can fail catastrophically. With our networks mapping back to the basics of the manifold, our research goes back to basics and pursues a self-contained paradigm for Transformer-based diffusion on raw natural data.

70HF ↗arXiv ↗
457

Variational Reasoning for Language Models

Xiangxin Zhou, Zichen Liu, Haonan Wang +5 authors

A variational reasoning framework treats thinking traces as latent variables, optimizing them through variational inference to improve language model reasoning.

70variational reasoning frameworklatent variablesHF ↗arXiv ↗
463

Competitive Programming with Large Reasoning Models

OpenAI, Ahmed El-Kishky, Alexander Wei +22 authors

General-purpose reinforcement learning applied to large language models outperforms domain-specific systems in complex coding and reasoning tasks, achieving top results in competitions without hand-crafted strategies.

69reinforcement learninglarge language modelsHF ↗arXiv ↗
464

Magistral

Mistral-AI, Abhinav Rastogi, Albert Q. Jiang +97 authors

Magistral, a scalable reinforcement learning pipeline, demonstrates that RL can enhance multimodal understanding and instruction following in large language models without requiring existing RL traces.

69reinforcement learningRLHF ↗arXiv ↗
467

Deep Researcher with Test-Time Diffusion

Rujun Han, Yanfei Chen, Zoey CuiZhu +15 authors

TTD-DR, a diffusion-based framework, generates high-quality research reports by iteratively refining a preliminary draft with external information and self-evolutionary algorithms, outperforming existing deep research agents.

69Large Language Models (LLMs)Test-Time Diffusion Deep Researcher (TTD-DR)HF ↗arXiv ↗
472

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

Jiaqi Tang, Jianmin Chen, Wei Wei +7 authors

A novel framework, Robust-R1, enhances multimodal large language models' robustness to visual degradations through explicit modeling, supervised fine-tuning, reward-driven alignment, and dynamic reasoning depth scaling, achieving state-of-the-art performance on real-world degradation benchmarks.

68multimodal large language modelsvisual degradationsHF ↗arXiv ↗
477

GHOST 2.0: generative high-fidelity one shot transfer of heads

Alexander Groshev, Anastasiia Iashchenko, Pavel Paramonov +2 authors

GHOST 2.0, consisting of an Aligner and Blender module, achieves state-of-the-art results in head swapping by preserving identity information, handling extreme poses, and seamlessly integrating the reenacted head into the target background.

67AlignerBlenderHF ↗arXiv ↗
478

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Prateek Chhikara, Dev Khant, Saket Aryan +2 authors

Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.

67Mem0memory-centric architectureHF ↗arXiv ↗
480

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Omkar Thawakar, Dinura Dissanayake, Ketan More +12 authors

A framework for evaluating and improving step-by-step visual reasoning in large language models using a specialized benchmark and a novel multimodal model trained with curriculum learning.

67visual reasoninglarge language modelsHF ↗arXiv ↗
16 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号