TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 10 – Nov 16, 2025
本周最热218

Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds

Weihao Tan, Xiangyang Li, Yunhao Fang +11 authors

Lumine, a vision-language model-based agent, completes complex missions in real-time across different 3D open-world environments with human-like efficiency and zero-shot cross-game generalization.

vision-language modelend-to-end3D open-world environmentshuman-like interactionHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Grounding Computer Use Agents on Human Demonstrations

Aarash Feizi, Shravan Nayak, Xiangru Jian +14 authors

GroundCUA, a large-scale desktop grounding dataset, enables the development of GroundNext models that achieve state-of-the-art performance in mapping instructions to UI elements with less training data.

107groundingnatural language instructionsHF ↗arXiv ↗
06

Depth Anything 3: Recovering the Visual Space from Any Views

Haotong Lin, Sili Chen, Junhao Liew +5 authors

Depth Anything 3 (DA3) uses a plain transformer for geometry prediction from visual inputs, achieving state-of-the-art results in camera pose estimation, any-view geometry, visual rendering, and monocular depth estimation.

103plain transformervanilla DINO encoderHF ↗arXiv ↗
09

IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction

Guoxin Chen, Zile Qiao, Xuanzhong Chen +13 authors

IterResearch, an iterative deep-research paradigm, improves long-horizon reasoning by reformulating it as a Markov Decision Process with strategic workspace reconstruction and Efficiency-Aware Policy Optimization, achieving better performance and interaction scaling compared to existing agents.

80deep-research agentsdynamic reasoningHF ↗arXiv ↗
11

MADD: Multi-Agent Drug Discovery Orchestra

Gleb V. Solovev, Alina B. Zhidkovskaya, Anastasia Orlova +18 authors

MADD, a multi-agent system integrating large language models and specialized models, streamlines hit identification in early drug discovery with superior performance and accessibility.

57large language modelsmulti-agent systemsHF ↗arXiv ↗
13

Black-Box On-Policy Distillation of Large Language Models

Tianzhu Ye, Li Dong, Zewen Chi +3 authors

Generative Adversarial Distillation (GAD) enhances black-box distillation by framing the student model as a generator and using a discriminator to provide adaptive feedback, surpassing traditional sequence-level knowledge distillation.

54black-box distillationlarge language models (LLMs)HF ↗arXiv ↗
15

Visual Spatial Tuning

Rui Yang, Ziyu Zhu, Yanwei Li +9 authors

A framework called Visual Spatial Tuning (VST) enhances the spatial abilities of Vision-Language Models (VLMs) through progressive training with specialized datasets, achieving state-of-the-art results on spatial benchmarks.

53Visual Spatial TuningVSTHF ↗arXiv ↗
16

DeepEyesV2: Toward Agentic Multimodal Model

Jack Hong, Chenxiao Zhao, ChengLin Zhu +3 authors

DeepEyesV2, an agentic multimodal model, uses a two-stage training pipeline to effectively integrate tool use, demonstrating robust performance across real-world reasoning tasks.

47DeepEyesV2reinforcement learningHF ↗arXiv ↗
17

Adaptive Multi-Agent Response Refinement in Conversational Systems

Soyeong Jeong, Aparna Elangovan, Emine Yilmaz +1 authors

A multi-agent framework enhances conversational quality by refining responses through agents responsible for factuality, personalization, and coherence, outperforming existing methods on challenging datasets.

42Large Language Models (LLMs)conversational systemsHF ↗arXiv ↗
18

Motif 2 12.7B technical report

Junghwan Lim, Sungmin Lee, Dongseok Kim +22 authors

Motif-2-12.7B combines architectural innovations and system optimizations to enhance efficiency and performance in large language models.

41Grouped Differential Attention (GDA)PolyNorm activationsHF ↗arXiv ↗
20

The Path Not Taken: RLVR Provably Learns Off the Principals

Hanqing Zhu, Zhenyu Zhang, Hanxian Huang +11 authors

Reinforcement Learning with Verifiable Rewards (RLVR) improves large language models by updating a limited set of parameters, which is explained by a Three-Gate Theory, revealing distinct optimization dynamics compared to supervised fine-tuning.

37Reinforcement Learning with Verifiable RewardsRLVRHF ↗arXiv ↗
22

KLASS: KL-Guided Fast Inference in Masked Diffusion Models

Seo Hyun Kim, Sunwoo Hong, Hojung Jung +2 authors

KL-Adaptive Stability Sampling (KLASS) accelerates diffusion-based generation by identifying stable predictions, achieving significant speedups and quality improvements across various domains.

37masked diffusion modelslanguage generationHF ↗arXiv ↗
26

Robot Learning from a Physical World Model

Jiageng Mao, Sicheng He, Hao-Ning Wu +9 authors

PhysWorld integrates video generation and physical world modeling to enable accurate robotic manipulation from visual demonstrations without real robot data.

32video generationphysical world modelingHF ↗arXiv ↗
27

Hail to the Thief: Exploring Attacks and Defenses in Decentralised GRPO

Nikolay Blagoev, Oğuzhan Ersoy, Lydia Yiyu Chen

The study identifies and defends against adversarial attacks in decentralized Group Relative Policy Optimization (GRPO) for Large Language Models (LLMs), demonstrating attack success rates of up to 100% and proposing effective defense mechanisms.

29Group Relative Policy OptimizationGRPOHF ↗arXiv ↗
30

VideoSSR: Video Self-Supervised Reinforcement Learning

Zefeng He, Xiaoye Qu, Yafu Li +3 authors

A novel video self-supervised reinforcement learning framework, VideoSSR, enhances MLLM performance across various video understanding tasks by leveraging intrinsic video information.

26Reinforcement Learning with Verifiable RewardsMultimodal Large Language ModelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号