TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 3 – Nov 9, 2025
本周最热242

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Jingqi Tong, Yurong Mou, Hangcheng Li +11 authors

The "Thinking with Video" paradigm enhances multimodal reasoning by integrating video generation models, demonstrated through the Video Thinking Benchmark and improved performance on both vision and text tasks.

Thinking with TextThinking with Imageslarge language modelsVision Language ModelsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Diffusion Language Models are Super Data Learners

Jinjie Ni, Qian Liu, Longxu Dou +5 authors

Diffusion language models outperform autoregressive models in low-data settings due to any-order modeling, iterative bidirectional denoising, and Monte Carlo augmentation, and maintain advantages even at scale.

132diffusion language modelsautoregressive modelsHF ↗arXiv ↗
05

V-Thinker: Interactive Thinking with Images

Runqi Qiao, Qiuna Tan, Minghan Yang +10 authors

V-Thinker, a multimodal reasoning assistant using reinforcement learning, enhances image-interactive thinking by synthesizing datasets and aligning perception for improved performance in vision-centric tasks.

98multimodal modelsimage interactionHF ↗arXiv ↗
08

Scaling Agent Learning via Experience Synthesis

Zhaorun Chen, Zhuokai Zhao, Kai Zhang +15 authors

DreamGym is a unified framework that synthesizes diverse experiences for scalable online RL training, improving agent performance and reducing real-world interactions.

83reinforcement learninglarge language modelHF ↗arXiv ↗
10

Continuous Autoregressive Language Models

Chenze Shao, Darren Li, Fandong Meng +1 authors

Continuous Autoregressive Language Models (CALM) improve language model efficiency by predicting continuous vectors instead of discrete tokens, reducing computational cost while maintaining performance.

75Continuous Autoregressive Language ModelsCALMHF ↗arXiv ↗
12

π_RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Kang Chen, Zhihao Liu, Tonghe Zhang +10 authors

The framework π<sub>RL</sub> uses reinforcement learning to train flow-based Vision-Language-Action models, addressing challenges with intractable action log-likelihoods and achieving significant performance improvements over supervised fine-tuning.

66reinforcement learningsupervised fine-tuningHF ↗arXiv ↗
17

World Simulation with Video Foundation Models for Physical AI

NVIDIA, Arslan Ali, Junjie Bai +86 authors

Cosmos-Predict2.5 and Cosmos-Transfer2.5 are advanced Physical AI models that unify text, image, and video generation, improve video quality and instruction alignment, and enable Sim2Real and Real2Real world translation with higher fidelity.

47flow-based architectureText2WorldHF ↗arXiv ↗
19

Cambrian-S: Towards Spatial Supersensing in Video

Shusheng Yang, Jihan Yang, Pinzhi Huang +12 authors

Progress in multimodal intelligence requires a shift to supersensing, including semantic perception, event cognition, spatial cognition, and predictive modeling, demonstrated through VSI-SUPER benchmarks and a self-supervised predictive sensing approach.

40supersensingsemantic perceptionHF ↗arXiv ↗
20

UniREditBench: A Unified Reasoning-based Image Editing Benchmark

Feng Han, Yibin Wang, Chenglin Li +8 authors

UniREditBench is a unified benchmark for reasoning-based image editing that addresses limitations in existing benchmarks by including multi-object interactions, game-world scenarios, and multimodal dual-reference evaluation.

39multi-modal generative modelsimage editingHF ↗arXiv ↗
23

MotionStream: Real-Time Video Generation with Interactive Motion Controls

Joonghyuk Shin, Zhengqi Li, Richard Zhang +4 authors

MotionStream enables real-time video generation with sub-second latency and up to 29 FPS by distilling a text-to-video model with motion control into a causal student using Self Forcing with Distribution Matching Distillation and sliding-window causal attention with attention sinks.

33motion-conditioned video generationMotionStreamHF ↗arXiv ↗
24

NVIDIA Nemotron Nano V2 VL

NVIDIA, Amala Sanjay Deshmukh, Kateryna Chumachenko +123 authors

Nemotron Nano V2 VL, a hybrid Mamba-Transformer LLM, improves document and video understanding through enhanced architecture and token reduction techniques.

32Mamba-Transformertoken reduction techniquesHF ↗arXiv ↗
26

Defeating the Training-Inference Mismatch via FP16

Penghui Qi, Zichen Liu, Xiangxin Zhou +4 authors

Using FP16 precision in reinforcement learning fine-tuning of large language models improves stability, convergence, and performance by addressing numerical mismatches.

32reinforcement learninglarge language modelsHF ↗arXiv ↗
30

Step-Audio-EditX Technical Report

Chao Yan, Boyong Wu, Peng Yang +10 authors

Step-Audio-EditX, an open-source LLM-based audio model, excels in expressive and iterative audio editing and zero-shot TTS using large-margin synthetic data.

30LLM-based audio modelexpressive audio editingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号