TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

September 2025

50 篇论文 · 按点赞排序

31

Video models are zero-shot learners and reasoners

Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6 authors

Veo 3, a generative video model, exhibits zero-shot capabilities across various visual tasks, suggesting a trajectory towards becoming a unified, generalist vision foundation model.

101Large Language ModelsLLMsHF ↗arXiv ↗
32

From Editor to Dense Geometry Estimator

JiYuan Wang, Chunyu Lin, Lei Sun +6 authors

FE2E, a framework using a Diffusion Transformer for dense geometry prediction, outperforms generative models in zero-shot monocular depth and normal estimation with improved performance and efficiency.

96text-to-imagedense predictionHF ↗arXiv ↗
34

Tree Search for LLM Agent Reinforcement Learning

Yuxiang Ji, Ziyu Ma, Yong Wang +3 authors

Tree-based Group Relative Policy Optimization (Tree-GRPO) enhances reinforcement learning for large language models by using tree search to improve rollouts and estimate grouped relative advantages, outperforming chain-based methods.

92reinforcement learninglarge language modelsHF ↗arXiv ↗
36

Seedream 4.0: Toward Next-generation Multimodal Image Generation

Team Seedream, Yunpeng Chen, Yu Gao +47 authors

Seedream 4.0 is a high-performance multimodal image generation system that integrates text-to-image synthesis, image editing, and multi-image composition using a diffusion transformer and VAE, achieving state-of-the-art results with efficient training and inference.

89diffusion transformerVAEHF ↗arXiv ↗
42

VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

Dongfu Jiang, Yi Lu, Zhuofeng Li +9 authors

VerlTool is a unified and modular framework for Agentic Reinforcement Learning with Tool use, addressing inefficiencies in existing approaches and providing competitive performance across multiple domains.

82Reinforcement Learning with Verifiable RewardsAgentic Reinforcement Learning with Tool useHF ↗arXiv ↗
45

Towards a Unified View of Large Language Model Post-Training

Xingtai Lv, Yuxin Zuo, Youbang Sun +9 authors

A unified policy gradient estimator and Hybrid Post-Training algorithm effectively combine online and offline data for post-training language models, improving performance across various benchmarks.

77Reinforcement LearningSupervised Fine-TuningHF ↗arXiv ↗
47

Open Data Synthesis For Deep Research

Ziyi Xia, Kun Luo, Hongjin Qian +1 authors

InfoSeek is a scalable framework for generating complex Deep Research tasks by synthesizing hierarchical constraint satisfaction problems, enabling models to outperform larger baselines on challenging benchmarks.

74Hierarchical Constraint Satisfaction ProblemsHCSPsHF ↗arXiv ↗
49

RewardDance: Reward Scaling in Visual Generation

Jie Wu, Yu Gao, Zilyu Ye +9 authors

RewardDance is a scalable reward modeling framework that aligns with VLM architectures, enabling effective scaling of RMs and resolving reward hacking issues in generation models.

73CLIP-based RMsBradley-Terry lossesHF ↗arXiv ↗
50

Variational Reasoning for Language Models

Xiangxin Zhou, Zichen Liu, Haonan Wang +5 authors

A variational reasoning framework treats thinking traces as latent variables, optimizing them through variational inference to improve language model reasoning.

70variational reasoning frameworklatent variablesHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号