TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 27 – Feb 2, 2025
本周最热128

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Tianzhe Chu, Yuexiang Zhai, Jihan Yang +6 authors

Reinforcement learning demonstrates superior generalization compared to supervised fine-tuning across textual and visual domains, while supervised fine-tuning stabilizes the model output format essential for RL performance.

supervised fine-tuningreinforcement learninggeneralizationmemorizationHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1155 authors

HLE is a challenging multi-modal benchmark that highlights the limitations of current LLMs in closed-ended academic questions.

78large language model (LLM)benchmarksHF ↗arXiv ↗
04

Qwen2.5-1M Technical Report

An Yang, Bowen Yu, Chengyuan Li +25 authors

The Qwen2.5-1M series models extend context length to 1 million tokens with enhanced long-context capabilities, employing techniques like long data synthesis and progressive pre-training, and are supported by an open-source inference framework with sparse attention and kernel optimizations.

72long-context pre-traininglong data synthesisHF ↗arXiv ↗
05

Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang +90 authors

Baichuan-Omni-1.5 is an omni-modal model with end-to-end audio generation, featuring a comprehensive data pipeline, audio-tokenizer, and multi-stage training strategy for superior performance across multimodal tasks.

62omni-modal modelaudio-tokenizerHF ↗arXiv ↗
08

Chain-of-Retrieval Augmented Generation

Liang Wang, Haonan Chen, Nan Yang +3 authors

CoRAG, a multi-step retrieval and reasoning approach, enhances RAG models by dynamically refining queries and using rejection sampling to improve performance, especially in multi-hop question answering.

58RAG modelsCoRAGHF ↗arXiv ↗
10

Atla Selene Mini: A General Purpose Evaluation Model

Andrei Alexandru, Antonia Calvi, Henry Broomfield +9 authors

Atla Selene Mini, an 8B language model-as-a-judge, excels across various benchmarks using enhanced data curation and a combined training approach, achieving top performance in zero-shot evaluations and real-world scenarios.

35data curationsynthetically generated critiquesHF ↗arXiv ↗
12

Towards General-Purpose Model-Free Reinforcement Learning

Scott Fujimoto, Pierluca D'Oro, Amy Zhang +2 authors

The paper presents MR.Q, a model-free deep RL algorithm that uses model-based representations to improve performance across diverse benchmarks without incurring high computational costs.

31reinforcement learningmodel-based RLHF ↗arXiv ↗
14

RL + Transformer = A General-Purpose Problem Solver

Micah Rentschler, Jesse Roberts

A pre-trained transformer fine-tuned with reinforcement learning develops the ability to solve unseen problems with sample efficiency and adaptability, demonstrating robust meta-learning.

28transformerreinforcement learningHF ↗arXiv ↗
15

Redundancy Principles for MLLMs Benchmarks

Zicheng Zhang, Xiangyu Zhao, Xinyu Fang +6 authors

With the rapid iteration of Multi-modality Large Language Models (MLLMs) and the evolving demands of the field, the number of benchmarks produced annually has surged into the hundreds. The rapid growth has inevitably led to significant redundancy among benchmarks. Therefore, it is crucial to take a step back and critically assess the current state of redundancy and propose targeted principles for constructing effective MLLM benchmarks. In this paper, we focus on redundancy from three key perspectives: 1) Redundancy of benchmark capability dimensions, 2) Redundancy in the number of test questions, and 3) Cross-benchmark redundancy within specific domains. Through the comprehensive analysis over hundreds of MLLMs' performance across more than 20 benchmarks, we aim to quantitatively measure the level of redundancy lies in existing MLLM evaluations, provide valuable insights to guide the future development of MLLM benchmarks, and offer strategies to refine and address redundancy issues effectively.

28sequence-to-sequenceHF ↗arXiv ↗
23

Open Problems in Mechanistic Interpretability

Lee Sharkey, Bilal Chughtai, Joshua Batson +26 authors

Mechanistic interpretability aims to understand neural network mechanisms to enhance AI assurance and address scientific questions about intelligence, highlighting open problems in methods, application, and socio-technical challenges.

21mechanistic interpretabilityneural networksHF ↗arXiv ↗
28

People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text

Jenna Russell, Marzena Karpinska, Mohit Iyyer

In this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. Our experiments show that annotators who frequently use LLMs for writing tasks excel at detecting AI-generated text, even without any specialized training or feedback. In fact, the majority vote among five such "expert" annotators misclassifies only 1 of 300 articles, significantly outperforming most commercial and open-source detectors we evaluated even in the presence of evasion tactics like paraphrasing and humanization. Qualitative analysis of the experts' free-form explanations shows that while they rely heavily on specific lexical clues ('AI vocabulary'), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity) that are challenging to assess for automatic detectors. We release our annotated dataset and code to spur future research into both human and automated detection of AI-generated text.

16LLMsGPT-4HF ↗arXiv ↗
29

Exploring the sustainable scaling of AI dilemma: A projective study of corporations' AI environmental impacts

Clément Desroches, Martin Chauvin, Louis Ladan +3 authors

The paper presents a methodology to estimate the environmental impact of a company's AI portfolio, highlighting the significant energy consumption of large generative AI models and advocating for coordinated efforts and standardized frameworks to align AI development with net-zero goals.

16Large Language Models (LLMs)environmental impactHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号