TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 18 – Aug 24, 2025
本周最热317

DINOv3

Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23 authors

DINOv3, a self-supervised learning model, achieves superior performance across various vision tasks by scaling datasets and models, addressing dense feature degradation, and enhancing flexibility with post-hoc strategies.

self-supervised learningDINOv3data preparationdesignHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Intern-S1: A Scientific Multimodal Foundation Model

Lei Bai, Zhongrui Cai, Maosong Cao +172 authors

Intern-S1, a multimodal Mixture-of-Experts model with extensive pre-training and reinforcement learning, achieves top-tier performance in general reasoning and outperforms closed-source models in scientific tasks.

274Mixture-of-Experts (MoE)reinforcement learning (RL)HF ↗arXiv ↗
04

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia +39 authors

Ovis2.5, a native-resolution vision transformer with multimodal reasoning, achieves state-of-the-art performance on various benchmarks through advanced training techniques and efficient scaling methods.

116vision transformernative-resolutionHF ↗arXiv ↗
05

SSRL: Self-Search Reinforcement Learning

Yuchen Fan, Kaiyan Zhang, Heng Zhou +15 authors

LLMs can serve as efficient simulators for RL tasks by leveraging internal knowledge, reducing reliance on external search engines and improving sim-to-real transfer.

97large language modelsLLMsHF ↗arXiv ↗
06

Deep Think with Confidence

Yichao Fu, Xuewei Wang, Yuandong Tian +1 authors

DeepConf enhances reasoning efficiency and performance by filtering low-quality reasoning traces using model-internal confidence signals, achieving high accuracy and reducing token generation.

92Deep Think with ConfidenceDeepConfHF ↗arXiv ↗
08

Thyme: Think Beyond Images

Yi-Fan Zhang, Xingyu Lu, Shukang Yin +17 authors

Thyme, a novel paradigm, enables MLLMs to autonomously perform image manipulations and computations, enhancing performance in perception and reasoning tasks through a two-stage training strategy and GRPO-ATS algorithm.

81MLLMsthink with imagesHF ↗arXiv ↗
12

Mobile-Agent-v3: Foundamental Agents for GUI Automation

Jiabo Ye, Xi Zhang, Haiyang Xu +12 authors

GUI-Owl and Mobile-Agent-v3 are open-source GUI agent models and frameworks that achieve state-of-the-art performance across various benchmarks using innovations in environment infrastructure, agent capabilities, and scalable reinforcement learning.

66GUI agent modelSelf-Evolving GUI Trajectory ProductionHF ↗arXiv ↗
13

4DNeX: Feed-Forward 4D Generative Modeling Made Easy

Zhaoxi Chen, Tianqi Liu, Long Zhuo +6 authors

4DNeX generates high-quality dynamic 3D scene representations from a single image using a fine-tuned pretrained video diffusion model, outperforming existing methods in efficiency and generalizability.

62feed-forward framework4D scene representationsHF ↗arXiv ↗
19

Prompt Orchestration Markup Language

Yuge Zhang, Nan Chen, Jiahang Xu +1 authors

POML addresses challenges in prompting Large Language Models by providing a structured, data-integrated, and format-sensitive markup language with templating and developer tools.

49POMLPrompt Orchestration Markup LanguageHF ↗arXiv ↗
20

Next Visual Granularity Generation

Yikai Wang, Zhouxia Wang, Zhonghua Wu +3 authors

A novel Next Visual Granularity (NVG) framework generates images by iteratively refining a sequence of visual granularities, outperforming existing methods in class-conditional image generation.

49Next Visual Granularity (NVG)visual granularity sequenceHF ↗arXiv ↗
27

Waver: Wave Your Way to Lifelike Video Generation

Yifu Zhang, Hao Yang, Yuqi Zhang +7 authors

Waver, a high-performance foundation model, generates high-quality videos and images using a Hybrid Stream DiT architecture, excelling in text-to-video, image-to-video, and text-to-image tasks.

39Hybrid Stream DiTmodality alignmentHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号