TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 17 – Aug 23, 2026

50 篇论文 · 按点赞排序

32

Agent Lightning v1.0: Towards Harnessed Agentic RL

Zhiyuan He, Siwei Zhang, Zhiwen Zhou +7 authors

Agent Lightning v1.0 enables reproducible reinforcement learning for arbitrary agent harnesses, substantially improving coding-agent performance with minimal data and compute.

35agent harnessesharnessed agentic RLHF ↗arXiv ↗
34

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Mengru Wang, Haozhe Luo, Zhenqian Xu +6 authors

Retrieved memories can induce reasoning errors and belief distortions in large language models, and an inference-time strategy helps avoid these cognitive traps while maintaining benchmark performance.

34memory-induced cognitive trapsReasoning FixationHF ↗arXiv ↗
42

GenRouter: Unified Workflow Routing for Agentic Image Generation

Harold Haodong Chen, Zhiyu Hou, Wen-Jie Shu +4 authors

GenRouter is a unified routing framework that adaptively directs prompts to optimal agentic image-generation workflows, cutting costs and latency while improving visual alignment and enabling continuous self-evolution.

28agentic image generationworkflow routingHF ↗arXiv ↗
44

MobileMem: Learning from a Year of Mobile Experiences

Xinle Deng, Yida Xue, Xiangyuan Ru +14 authors

MobileMem is a benchmark and framework for evaluating on-device long-term memory through year-scale, multimodal mobile experience trajectories that require temporal reasoning, knowledge updating, and preference inference.

26on-device long-term memorymultimodalHF ↗arXiv ↗
45

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Yiming Du, Yuxin Jiang, Tao Yuan +9 authors

LEGO-RL connects native coding-agent harnesses to scalable policy-gradient training via in-process LLM proxying, sandbox orchestration, and integrated monitoring, improving sparse MoE model performance across multiple harnesses.

25reinforcement learningpolicy-gradient optimizationHF ↗arXiv ↗
46

Latent On-Policy Self-Distillation

Guibin Zhang, Jiayang Lyu, Ran Sun +4 authors

Latent On-Policy Self-Distillation learns privileged teaching context end-to-end from experience to provide dense token-level supervision, improving agent performance and efficiency.

25on-policy self-distillationprivileged self-teacherHF ↗arXiv ↗
49

Repo0: Design-Driven Zero-to-All Code Generation

Silin Chen, Haoyi Teng, Xiaodong Gu +5 authors

Repo0 uses a dual-graph architectural state and modularity-guided structural evolution to generate complete software repositories from natural-language requirements with high functionality coverage.

22zero-to-all code generationDual-DAGHF ↗arXiv ↗
50

Looped Language Models Improve Compositional Tool Calling

Andrei Cristian Popescu, Haitz Sáez de Ocáriz Borde, Pietro Liò

Looped language models improve compositional, multi-step tool use through recurrent computation, with adaptive inference balancing accuracy and compute cost.

22looped language modelscompositional tool-callingHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号