TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 23 – Mar 29, 2026
本周最热139

MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding

Hejun Dong, Junbo Niu, Bin Wang +3 authors

MinerU-Diffusion is a diffusion-based framework that replaces autoregressive decoding with parallel diffusion denoising for document OCR, improving robustness and decoding speed.

diffusion-based frameworkautoregressive decodingparallel diffusion denoisingblock-wise diffusion decoderHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

Yicheng Zou, Dongsheng Zhu, Lin Zhu +171 authors

Intern-S1-Pro is a one-trillion-parameter scientific multimodal foundation model that enhances general and scientific capabilities through advanced agent functionalities and specialized task mastery across multiple scientific disciplines.

134multimodal foundation modelreinforcement learningHF ↗arXiv ↗
05

PixelSmile: Toward Fine-Grained Facial Expression Editing

Jiabin Hua, Hengyuan Xu, Aojie Li +4 authors

A diffusion framework called PixelSmile is introduced that disentangles facial expression semantics through symmetric joint training and contrastive learning to enable precise, controllable, and fine-grained expression editing with robust identity preservation.

118diffusion frameworkfacial expression editingHF ↗arXiv ↗
14

Voxtral TTS

Alexander H. Liu, Alexis Tacnet, Andy Ehrenberg +184 authors

Voxtral TTS is a multilingual text-to-speech model that generates natural speech from short reference audio using a hybrid architecture combining semantic token generation and flow-matching for acoustic tokens.

63text-to-speechauto-regressive generationHF ↗arXiv ↗
20

DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models

Jaewon Min, Jaeeun Lee, Yeji Choi +7 authors

Optical flow models trained on high-quality data often degrade severely when confronted with real-world corruptions such as blur, noise, and compression artifacts. To overcome this limitation, we formulate Degradation-Aware Optical Flow, a new task targeting accurate dense correspondence estimation from real-world corrupted videos. Our key insight is that the intermediate representations of image restoration diffusion models are inherently corruption-aware but lack temporal awareness. To address this limitation, we lift the model to attend across adjacent frames via full spatio-temporal attention, and empirically demonstrate that the resulting features exhibit zero-shot correspondence capabilities. Based on this finding, we present DA-Flow, a hybrid architecture that fuses these diffusion features with convolutional features within an iterative refinement framework. DA-Flow substantially outperforms existing optical flow methods under severe degradation across multiple benchmarks.

52optical flowimage restorationHF ↗arXiv ↗
21

Hyperagents

Jenny Zhang, Bingchen Zhao, Wannan Yang +5 authors

Hyperagents represent a self-referential framework that integrates task and meta-agents into a single editable program, enabling metacognitive self-modification and open-ended improvement across diverse computational domains.

51self-improving AI systemsDarwin Gödel MachineHF ↗arXiv ↗
22

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

Yan Shu, Bin Ren, Zhitong Xiong +4 authors

TerraScope is a unified vision-language model that enables pixel-grounded geospatial reasoning through modality-flexible and multi-temporal capabilities, evaluated on a new benchmark with detailed visual reasoning outputs.

51vision-language modelspixel-grounded reasoningHF ↗arXiv ↗
24

Repurposing Geometric Foundation Models for Multi-view Diffusion

Wooseok Jang, Seonghu Jeon, Jisang Han +5 authors

Geometric Latent Diffusion (GLD) framework utilizes geometric foundation models' feature space as latent space for novel view synthesis, achieving superior 2D and 3D performance while reducing training time significantly.

50generative latent spacesnovel view synthesisHF ↗arXiv ↗
27

EVA: Efficient Reinforcement Learning for End-to-End Video Agent

Yaolun Zhang, Ruohui Wang, Jiahao Wang +6 authors

EVA is an efficient reinforcement learning framework for video understanding that enables adaptive reasoning through iterative planning and attention mechanisms, outperforming existing methods on multiple video benchmarks.

44multimodal large language modelsreinforcement learningHF ↗arXiv ↗
29

ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models

Thomas De Min, Subhankar Roy, Stéphane Lathuilière +2 authors

MLLMs demonstrate limited proactive behavior in requesting user interventions for challenging tasks, with performance hindered by conversational context and in-context learning biases, though reinforcement learning fine-tuning shows potential for learning such behaviors.

41MLLMsproactive behaviorHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号