TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 16 – Jun 22, 2025
本周最热279

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

MiniMax, Aili Chen, Aonian Li +124 authors

A hybrid-attention reasoning model called MiniMax-M1, featuring a Mixture-of-Experts architecture and lightning attention mechanism, is introduced for efficient long-input processing and reinforcement learning.

Mixture-of-Experts (MoE)lightning attention mechanismreinforcement learning (RL)CISPOHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Sekai: A Video Dataset towards World Exploration

Zhen Li, Chuanhao Li, Xiaofeng Mao +17 authors

Sekai, a worldwide video dataset with comprehensive annotations, is introduced to support world exploration applications, enhancing video generation models.

67first-person viewworldwide video datasetHF ↗arXiv ↗
06

Scaling Test-time Compute for LLM Agents

King Zhu, Hanhao Li, Siwei Wu +12 authors

Systematic exploration of test-time scaling methods in large language agents reveals that computational scaling improves performance, especially through parallel sampling, sequential revision, effective verification, and increased rollout diversity.

64parallel sampling algorithmssequential revision strategiesHF ↗arXiv ↗
11

Essential-Web v1.0: 24T tokens of organized web data

Essential AI, Andrew Hojel, Michael Pust +21 authors

A large, 24-trillion-token Essential-Web v1.0 dataset annotated with a multi-category taxonomy outperforms or is competitive with existing datasets in various domains using simple filtering techniques.

48Essential-WebEAI-Distill-0.5bHF ↗arXiv ↗
13

LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs

Xiaoran Liu, Zhigeng Liu, Zengfeng Huang +3 authors

This study investigates long-context performance of diffusion LLMs compared to auto-regressive LLMs, identifies their unique characteristics, and proposes LongLLaDA, a training-free method for extending context windows.

44diffusion LLMsauto-regressive LLMsHF ↗arXiv ↗
15

Discrete Diffusion in Large Language and Multimodal Models: A Survey

Runpeng Yu, Qi Li, Xinchao Wang

Discrete Diffusion Language Models (dLLMs) and Discrete Diffusion Multimodal Language Models (dMLLMs) enable parallel generation and faster inference compared to autoregressive models through denoising-based strategies and full attention mechanisms.

44Discrete Diffusion Language ModelsDiscrete Diffusion Multimodal Language ModelsHF ↗arXiv ↗
17

All is Not Lost: LLM Recovery without Checkpoints

Nikolay Blagoev, Oğuzhan Ersoy, Lydia Yiyu Chen

A novel method, CheckFree, and its extended version CheckFree+, efficiently recover from node failures during LLM training by substituting failed stages with averaged neighboring stages or through out-of-order pipeline execution, improving convergence time over existing checkpointing methods.

41LLMsdecentralized computationHF ↗arXiv ↗
20

Effective Red-Teaming of Policy-Adherent Agents

Itay Nakash, George Kour, Koren Lazar +3 authors

CRAFT, a multi-agent system using policy-aware persuasive strategies, challenges policy-adherent LLM-based agents in customer service to assess and improve their robustness against adversarial attacks.

39LLM-based agentspolicy-adherenceHF ↗arXiv ↗
21

The Diffusion Duality

Subham Sekhar Sahoo, Justin Deschenaux, Aaron Gokaslan +3 authors

Duo improves uniform-state discrete diffusion models by transferring techniques from Gaussian diffusion, enhancing training speed and enabling fast few-step text generation.

37discrete diffusion modelsGaussian diffusionHF ↗arXiv ↗
24

Efficient Medical VIE via Reinforcement Learning

Lijun Liu, Ruiyang Li, Zhaocheng Liu +5 authors

An RLVR framework using fine-tuned Qwen2.5-VL-7B achieves state-of-the-art performance in medical VIE with limited annotated samples, enhancing reasoning and balance between precision and recall.

31Reinforcement Learning with Verifiable Rewards (RLVR)JSON generationHF ↗arXiv ↗
25

Show-o2: Improved Native Unified Multimodal Models

Jinheng Xie, Zhenheng Yang, Mike Zheng Shou

Show-o2 leverages autoregressive modeling and flow matching within a 3D causal variational autoencoder to create unified visual representations for multimodal understanding and generation tasks.

30autoregressive modelingflow matchingHF ↗arXiv ↗
26

Reasoning with Exploration: An Entropy Perspective

Daixuan Cheng, Shaohan Huang, Xuekai Zhu +4 authors

Introducing an entropy-based term to the advantage function in reinforcement learning enhances exploratory reasoning in language models, leading to improved performance on complex reasoning tasks.

30reinforcement learningentropyHF ↗arXiv ↗
30

Test3R: Learning to Reconstruct 3D at Test Time

Yuheng Yuan, Qiuhong Shen, Shizun Wang +2 authors

Test3R, a test-time learning technique for 3D reconstruction, enhances geometric accuracy by optimizing network consistency using self-supervised learning on image triplets.

27DUSt3Rdense matchingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号