TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 24 – Nov 30, 2025

50 篇论文 · 按点赞排序

35

Pillar-0: A New Frontier for Radiology Foundation Models

Kumar Krishna Agrawal, Longchao Liu, Long Lian +11 authors

Pillar-0, a radiology foundation model pretrained on diverse imaging datasets, outperforms existing models across various tasks and extends to new applications using RATE for label extraction.

25foundation modelsvolumetric CTHF ↗arXiv ↗
36

NVIDIA Nemotron Parse 1.1

Kateryna Chumachenko, Amala Sanjay Deshmukh, Jarno Seppanen +30 authors

Nemotron-Parse-1.1 is a lightweight OCR and document parsing model with improved capabilities in general OCR, markdown formatting, structured table parsing, and text extraction from images, using an encoder-decoder architecture.

24OCRdocument parsingHF ↗arXiv ↗
37

WorldGen: From Text to Traversable and Interactive 3D Worlds

Dilin Wang, Hyunyoung Jung, Tom Monnier +22 authors

WorldGen converts text prompts into interactive 3D environments using LLM-driven reasoning, procedural generation, diffusion-based 3D generation, and object-aware decomposition, enabling creators to build coherent, navigable worlds efficiently.

24LLM-driven scene layout reasoningprocedural generationHF ↗arXiv ↗
39

HunyuanOCR Technical Report

Hunyuan Vision Team, Pengyuan Lyu, Xingyu Wan +23 authors

HunyuanOCR, a lightweight Vision-Language Model, achieves state-of-the-art performance in OCR tasks through a unified end-to-end architecture combining Vision Transformer and lightweight LLM, supported by data-driven and RL strategies.

23Vision-Language ModelVision TransformerHF ↗arXiv ↗
40

Fara-7B: An Efficient Agentic Model for Computer Use

Ahmed Awadallah, Yash Lara, Raghav Magazine +9 authors

FaraGen creates synthetic datasets for computer use agents, enabling the training of efficient and high-performing models like Fara-7B on diverse web tasks, outperforming larger models on benchmarks.

22FaraGensynthetic data generationHF ↗arXiv ↗
43

Loomis Painter: Reconstructing the Painting Process

Markus Pobitzer, Chang Liu, Chenyi Zhuang +3 authors

A unified framework using diffusion models with semantic control and cross-medium style augmentation generates consistent and high-fidelity multi-media painting processes, supported by a comprehensive dataset and evaluation metrics.

20diffusion modelssemantics-driven style controlHF ↗arXiv ↗
46

Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO

Haoyang Hong, Jiajun Yin, Yuan Wang +14 authors

M-GRPO, an extension of Group Relative Policy Optimization for hierarchical multi-agent systems, improves stability and efficiency in tool-augmented reasoning tasks by aligning heterogeneous trajectories and decoupling agent training.

19multi-agent systemslarge language model (LLM)HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号