TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 1 – Sep 7, 2025
本周最热350

A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code

Keke Lian, Bin Wang, Lei Zhang +18 authors

A.S.E is a repository-level benchmark for evaluating the security of AI-generated code, highlighting challenges in secure coding and the limitations of LLMs in real-world scenarios.

large language modelsLLMsAI code generationsecurity evaluationHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu +22 authors

Agentic reinforcement learning transforms large language models into autonomous decision-making agents by leveraging temporally extended POMDPs, enhancing capabilities like planning and reasoning through reinforcement learning.

238agentic reinforcement learningLLM RLHF ↗arXiv ↗
07

From Editor to Dense Geometry Estimator

JiYuan Wang, Chunyu Lin, Lei Sun +6 authors

FE2E, a framework using a Diffusion Transformer for dense geometry prediction, outperforms generative models in zero-shot monocular depth and normal estimation with improved performance and efficiency.

96text-to-imagedense predictionHF ↗arXiv ↗
10

VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

Dongfu Jiang, Yi Lu, Zhuofeng Li +9 authors

VerlTool is a unified and modular framework for Agentic Reinforcement Learning with Tool use, addressing inefficiencies in existing approaches and providing competitive performance across multiple domains.

82Reinforcement Learning with Verifiable RewardsAgentic Reinforcement Learning with Tool useHF ↗arXiv ↗
12

Towards a Unified View of Large Language Model Post-Training

Xingtai Lv, Yuxin Zuo, Youbang Sun +9 authors

A unified policy gradient estimator and Hybrid Post-Training algorithm effectively combine online and offline data for post-training language models, improving performance across various benchmarks.

77Reinforcement LearningSupervised Fine-TuningHF ↗arXiv ↗
14

Open Data Synthesis For Deep Research

Ziyi Xia, Kun Luo, Hongjin Qian +1 authors

InfoSeek is a scalable framework for generating complex Deep Research tasks by synthesizing hierarchical constraint satisfaction problems, enabling models to outperform larger baselines on challenging benchmarks.

74Hierarchical Constraint Satisfaction ProblemsHCSPsHF ↗arXiv ↗
19

Robix: A Unified Model for Robot Interaction, Reasoning and Planning

Huang Fang, Mengxi Zhang, Heng Dong +6 authors

Robix, a unified vision-language model, integrates robot reasoning, task planning, and natural language interaction, demonstrating superior performance in interactive task execution through chain-of-thought reasoning and a three-stage training strategy.

53chain-of-thought reasoningthree-stage training strategyHF ↗arXiv ↗
23

Kwai Keye-VL 1.5 Technical Report

Biao Yang, Bin Wen, Boyang Ding +57 authors

Keye-VL-1.5 enhances video understanding through a Slow-Fast encoding strategy, progressive pre-training, and post-training reasoning improvements, outperforming existing models on video tasks while maintaining general multimodal performance.

41Slow-Fast video encoding strategyprogressive four-stage pre-trainingHF ↗arXiv ↗
26

Transition Models: Rethinking the Generative Learning Objective

Zidong Wang, Yiyuan Zhang, Xiaoyu Yue +4 authors

A novel generative paradigm, Transition Models (TiM), addresses the trade-off between computational cost and output quality in generative modeling by using a continuous-time dynamics equation.

29iterative diffusion modelscontinuous-time dynamics equationHF ↗arXiv ↗
27

NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings

Or Shachar, Uri Katz, Yoav Goldberg +1 authors

NER Retriever uses internal representations from large language models to perform zero-shot named entity retrieval by embedding entity mentions and type descriptions into a shared semantic space, outperforming lexical and dense sentence-level retrieval methods.

29NER Retrieverzero-shot retrievalHF ↗arXiv ↗
30

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers

Xingyue Huang, Rishabh, Gregor Franke +43 authors

The Loong Project introduces a framework for generating and verifying synthetic data to improve reasoning capabilities in Large Language Models through Reinforcement Learning with Verifiable Reward.

26Large Language ModelsReinforcement Learning with Verifiable RewardHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号