TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 25 – Aug 31, 2025
本周最热226

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Weiyun Wang, Zhangwei Gao, Lixin Gu +58 authors

InternVL 3.5 introduces Cascade RL, ViR, and DvD to enhance reasoning, efficiency, and performance in multimodal models.

Cascade RLoffline RLonline RLVisual Resolution RouterHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

VibeVoice Technical Report

Zhiliang Peng, Jianwei Yu, Wenhui Wang +10 authors

VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.

179next-token diffusioncontinuous speech tokenizerHF ↗arXiv ↗
03

AgentFly: Fine-tuning LLM Agents without Fine-tuning LLMs

Huichi Zhou, Yihang Chen, Siyuan Guo +8 authors

A novel memory-augmented reinforcement learning paradigm enables adaptive LLM agents to continually learn without fine-tuning, using episodic memory and a neural case-selection policy.

162Large Language Model (LLM)memory-based online reinforcement learningHF ↗arXiv ↗
04

rStar2-Agent: Agentic Reasoning Technical Report

Ning Shang, Yifei Liu, Yi Zhu +12 authors

rStar2-Agent, a 14B math reasoning model trained with agentic reinforcement learning, achieves state-of-the-art performance by efficiently handling complex problem-solving with advanced cognitive behaviors and minimal computational resources.

121agentic reinforcement learningCoTHF ↗arXiv ↗
06

Beyond Transcription: Mechanistic Interpretability in ASR

Neta Glazer, Yael Segal-Feldman, Hilit Segev +6 authors

Interpretability methods like logit lens, linear probing, and activation patching are applied to ASR to uncover internal dynamics, repetition hallucinations, and semantic biases, enhancing model transparency and robustness.

92logit lenslinear probingHF ↗arXiv ↗
08

Self-Rewarding Vision-Language Model via Reasoning Decomposition

Zongxia Li, Wenhao Yu, Chengsong Huang +8 authors

Vision-SR1 uses reinforcement learning to enhance visual reasoning in vision-language models by decomposing the process into visual perception and language reasoning stages, improving accuracy and reducing hallucinations.

85vision-language modelsvisual hallucinationsHF ↗arXiv ↗
12

Hermes 4 Technical Report

Ryan Teknium, Roger Jin, Jai Suphavadeeprasit +6 authors

Hermes 4, a hybrid reasoning model, integrates structured multi-turn reasoning with broad instruction-following, evaluated across various benchmarks including math, coding, knowledge, comprehension, and alignment.

57hybrid reasoning modelsstructured reasoningHF ↗arXiv ↗
19

AWorld: Orchestrating the Training Recipe for Agentic AI

Chengyue Yu, Siyuan Lu, Chenyi Zhuang +14 authors

AWorld, an open-source system for large-scale agent-environment interaction, accelerates experience collection and enhances reinforcement learning, leading to significant improvements in agentic AI performance on complex benchmarks.

39reinforcement learningQwen3-32BHF ↗arXiv ↗
20

MV-RAG: Retrieval Augmented Multiview Diffusion

Yosef Dayani, Omer Benishu, Sagie Benaim

MV-RAG enhances text-to-3D generation by retrieving 2D images and conditioning a multiview diffusion model to improve consistency and accuracy, especially for out-of-domain concepts.

38pretrained 2D diffusion priorsMV-RAGHF ↗arXiv ↗
24

Mixture of Contexts for Long Video Generation

Shengqu Cai, Ceyuan Yang, Lvmin Zhang +10 authors

Long video generation is addressed by introducing a sparse attention routing module, Mixture of Contexts, to efficiently manage long-term memory and retrieval in diffusion transformers.

35diffusion transformersself-attentionHF ↗arXiv ↗
27

Understanding Tool-Integrated Reasoning

Heng Lin, Zhongwen Xu

Tool-Integrated Reasoning (TIR) enhances Large Language Models (LLMs) by expanding their problem-solving capabilities through the use of external tools, and Advantage Shaping Policy Optimization (ASPO) improves model behavior and tool usage.

32Tool-Integrated ReasoningLarge Language ModelsHF ↗arXiv ↗
28

Spacer: Towards Engineered Scientific Inspiration

Minhyeong Lee, Suyoung Hwang, Seunghyun Moon +13 authors

Spacer, a scientific discovery system, uses deliberate decontextualization to generate creative and factually grounded scientific concepts from keyword sets, achieving high accuracy and similarity to leading publications.

32LLMsscientific discovery systemHF ↗arXiv ↗
29

Autoregressive Universal Video Segmentation Model

Miran Heo, Sukjun Hwang, Min-Hung Chen +4 authors

AUSM, an autoregressive universal segmentation model, unifies prompted and unprompted video segmentation by treating it as sequential mask prediction, achieving superior performance and faster training on standard benchmarks.

29sequential mask predictionlanguage modelingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号