TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

31

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu +106 authors

Step-Audio~2, an end-to-end multi-modal large language model, integrates latent audio encoding and reinforcement learning to achieve state-of-the-art performance in ASR, audio understanding, and speech conversation, incorporating discrete audio token generation and retrieval-augmented generation.

77latent audio encoderreasoning-centric reinforcement learningHF ↗arXiv ↗
33

Skywork-R1V3 Technical Report

Wei Shen, Jiangbo Pei, Yi Peng +7 authors

Skywork-R1V3, an open-source vision-language model, enhances visual reasoning through a post-training reinforcement learning framework, achieving state-of-the-art performance on multimodal reasoning tasks.

75vision-language modelvisual reasoningHF ↗arXiv ↗
37

Energy-Based Transformers are Scalable Learners and Thinkers

Alexi Gladstone, Ganesh Nanduru, Md Mofijul Islam +7 authors

Energy-Based Transformers (EBTs) improve model performance and scalability across modalities by learning to verify predictions through unsupervised learning and energy minimization.

71Energy-Based TransformersEnergy-Based ModelsHF ↗arXiv ↗
39

Deep Researcher with Test-Time Diffusion

Rujun Han, Yanfei Chen, Zoey CuiZhu +15 authors

TTD-DR, a diffusion-based framework, generates high-quality research reports by iteratively refining a preliminary draft with external information and self-evolutionary algorithms, outperforming existing deep research agents.

69Large Language Models (LLMs)Test-Time Diffusion Deep Researcher (TTD-DR)HF ↗arXiv ↗
41

π^3: Scalable Permutation-Equivariant Visual Geometry Learning

Yifan Wang, Jianjun Zhou, Haoyi Zhu +7 authors

A permutation-equivariant neural network, $\pi^3$, reconstructs visual geometry without a fixed reference view, achieving state-of-the-art performance in camera pose estimation, depth estimation, and point map reconstruction.

67feed-forward neural networkpermutation-equivariant architectureHF ↗arXiv ↗
43

BANG: Dividing 3D Assets via Generative Exploded Dynamics

Longwen Zhang, Qixuan Zhang, Haoran Jiang +4 authors

BANG is a generative approach using latent diffusion models and temporal attention to enable intuitive, part-level decomposition of 3D objects with precise control and multimodal interaction.

65latent diffusion modelexploded dynamicsHF ↗arXiv ↗
46

Ovis-U1 Technical Report

Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9 authors

Ovis-U1, a 3-billion-parameter unified model, integrates multimodal understanding, text-to-image generation, and image editing using a diffusion-based visual decoder and bidirectional token refiner, achieving state-of-the-art performance across various benchmarks.

63diffusion-based visual decoderbidirectional token refinerHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号