TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

394

World-in-World: World Models in a Closed-Loop World

Jiahan Zhang, Muqing Jiang, Nanru Dai +14 authors

World-in-World evaluates generative world models in closed-loop environments, emphasizing task success over visual quality and revealing insights into controllability, data scaling, and compute allocation.

78generative world modelsWMHF ↗arXiv ↗
395

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +661 authors

HLE is a challenging multi-modal benchmark that highlights the limitations of current LLMs in closed-ended academic questions.

78large language model (LLM)benchmarksHF ↗arXiv ↗
398

DDT: Decoupled Diffusion Transformer

Shuai Wang, Zhi Tian, Weilin Huang +1 authors

A decoupled diffusion transformer improves performance and training speed in image generation by separating semantic extraction and high-frequency decoding.

77diffusion transformersdenoising stepsHF ↗arXiv ↗
400

Towards a Unified View of Large Language Model Post-Training

Xingtai Lv, Yuxin Zuo, Youbang Sun +9 authors

A unified policy gradient estimator and Hybrid Post-Training algorithm effectively combine online and offline data for post-training language models, improving performance across various benchmarks.

77Reinforcement LearningSupervised Fine-TuningHF ↗arXiv ↗
404

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu +106 authors

Step-Audio~2, an end-to-end multi-modal large language model, integrates latent audio encoding and reinforcement learning to achieve state-of-the-art performance in ASR, audio understanding, and speech conversation, incorporating discrete audio token generation and retrieval-augmented generation.

76latent audio encoderreasoning-centric reinforcement learningHF ↗arXiv ↗
408

Skywork-R1V3 Technical Report

Wei Shen, Jiangbo Pei, Yi Peng +7 authors

Skywork-R1V3, an open-source vision-language model, enhances visual reasoning through a post-training reinforcement learning framework, achieving state-of-the-art performance on multimodal reasoning tasks.

75vision-language modelvisual reasoningHF ↗arXiv ↗
418

Open Data Synthesis For Deep Research

Ziyi Xia, Kun Luo, Hongjin Qian +1 authors

InfoSeek is a scalable framework for generating complex Deep Research tasks by synthesizing hierarchical constraint satisfaction problems, enabling models to outperform larger baselines on challenging benchmarks.

74Hierarchical Constraint Satisfaction ProblemsHCSPsHF ↗arXiv ↗
14 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号