TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 18 – May 24, 2026

50 篇论文 · 按点赞排序

32

Auditing Agent Harness Safety

Chengzhi Liu, Yichen Guo, Yepeng Liu +8 authors

LLM agents executing within execution harnesses can produce correct outputs while violating safety constraints during execution, necessitating trajectory-level auditing to ensure proper resource access and information flow across multi-agent systems.

56execution harnessestool dispatchingHF ↗arXiv ↗
34

Process Rewards with Learned Reliability

Jinyuan Li, Langlin Huang, Chengsong Huang +5 authors

BetaPRM introduces a distributional approach to process reward models that predicts both success probabilities and prediction reliability, enabling adaptive computation allocation that reduces token usage while maintaining accuracy.

53Process Reward ModelsBetaPRMHF ↗arXiv ↗
36

You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories

Zhepei Wei, Xinyu Zhu, Wei-Lin Chen +3 authors

Reinforcement learning with verifiable rewards parameter trajectories exhibit low-rank structures that enable efficient extrapolation through a simple linear regression method, demonstrating superior performance with reduced computational requirements.

51reinforcement learning with verifiable rewardsparameter trajectoriesHF ↗arXiv ↗
43

Toto 2.0: Time Series Forecasting Enters the Scaling Era

Emaad Khwaja, Chris Lettieri, Gerald Woo +10 authors

Time series foundation models demonstrate scalable forecasting performance across parameter sizes, with Toto 2.0 achieving state-of-the-art results on multiple benchmarks through a unified training approach.

40time series foundation modelsforecasting modelsHF ↗arXiv ↗
46

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning

Banghao Chi, Yining Xie, Mingyuan Wu +9 authors

Spreadsheet-RL is a reinforcement learning framework that trains specialized spreadsheet agents in realistic Excel environments, improving AI agent performance on both general and domain-specific spreadsheet tasks through automated data collection and domain-specific benchmarks.

36reinforcement learningfine-tuningHF ↗arXiv ↗
47

Harnessing LLM Agents with Skill Programs

Hongjun Liu, Yifei Ming, Shafiq Joty +1 authors

HASP introduces executable program functions that serve as active guardrails for LLM agents, enabling direct intervention in agent loops and improving performance across complex tasks.

36LLM agentsskill programsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号