TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 9 – Mar 15, 2026
本周最热211

Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning

Lei Huang, Xiang Cheng, Chenxiao Zhao +6 authors

Language feedback is leveraged in reinforcement learning to improve exploration efficiency and sample utilization through grouped critique aggregation and joint generation-refinement optimization.

reinforcement learningnatural language feedbackscalar rewardstargeted explorationHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

OpenClaw-RL: Train Any Agent Simply by Talking

Yinjie Wang, Xuyang Chen, Xiaolong Jin +2 authors

OpenClaw-RL framework enables policy learning from diverse next-state signals across multiple interaction modalities using asynchronous training with PRM judges and hindsight-guided distillation.

158agentic RLnext-state signalsHF ↗arXiv ↗
09

Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs

Zorik Gekhman, Roee Aharoni, Eran Ofek +3 authors

Reasoning in large language models enhances parametric knowledge recall through computational buffer and factual priming mechanisms, though it carries risks of hallucination that can be mitigated by prioritizing accurate reasoning trajectories.

76large language modelsparametric knowledge recallHF ↗arXiv ↗
12

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

Junyi Zhang, Charles Herrmann, Junhwa Hur +5 authors

LoGeR enables long-term 3D video reconstruction by combining bidirectional priors with a hybrid memory system that includes parametric Test-Time Training and non-parametric sliding window attention mechanisms.

63feedforward geometric foundation modelsquadratic attention complexityHF ↗arXiv ↗
13

How Far Can Unsupervised RLVR Scale LLM Training?

Bingxiang He, Yuxin Zuo, Zeyuan Liu +18 authors

Unsupervised reinforcement learning with verifiable rewards faces fundamental limitations in scaling large language model training due to inherent convergence properties and confidence-correction misalignment, though external methods show promise for overcoming these barriers.

59unsupervised reinforcement learningverifiable rewardsHF ↗arXiv ↗
17

LLM2Vec-Gen: Generative Embeddings from Large Language Models

Parishad BehnamGhader, Vaibhav Adlakha, Fabian David Schmidt +3 authors

LLM2Vec-Gen introduces a self-supervised method for text embedding that represents model responses through trainable special tokens, achieving superior performance on MTEB while reducing harmful content and improving reasoning capabilities.

44LLM-based text embedderscontrastive learningHF ↗arXiv ↗
18

Video-Based Reward Modeling for Computer-Use Agents

Linxin Song, Jieyu Zhang, Huanxin Sheng +6 authors

Video-execution reward modeling enables scalable evaluation of computer-using agents by predicting task success from user instructions and execution videos, outperforming proprietary models across multiple operating systems.

43computer-using agentsreward modelingHF ↗arXiv ↗
19

In-Context Reinforcement Learning for Tool Use in Large Language Models

Yaoqi Ye, Yiran Zhao, Keyu Duan +4 authors

In-Context Reinforcement Learning (ICRL) enables large language models to effectively use external tools through a reinforcement learning-only framework that eliminates the need for costly supervised fine-tuning by gradually reducing in-context examples during training.

43large language modelsreinforcement learningHF ↗arXiv ↗
20

Reasoning Models Struggle to Control their Chains of Thought

Chen Yueh-Han, Robert McCarthy, Bruce W. Lee +5 authors

Chain-of-thought controllability measures how effectively models can be constrained to follow reasoning steps, with findings showing significantly lower controllability in reasoning versus output generation, and varying impacts from model size, training methods, and task complexity.

41Chain-of-thought monitoringCoT controllabilityHF ↗arXiv ↗
22

Fish Audio S2 Technical Report

Shijia Liao, Yuxuan Wang, Songting Liu +11 authors

Fish Audio S2 is an open-source text-to-speech system with multi-speaker capabilities, multi-turn generation, and instruction-following control through natural-language descriptions, utilizing a multi-stage training approach and production-ready inference engine.

40text-to-speechmulti-speakerHF ↗arXiv ↗
24

Believe Your Model: Distribution-Guided Confidence Calibration

Xizhong Yang, Haotian Zhang, Huiming Wang +1 authors

Large reasoning models enhance prediction accuracy through test-time scaling techniques that generate multiple candidate responses, with the proposed DistriVoting method utilizing distributional priors and confidence scores to improve answer selection.

40Gaussian Mixture Modelsconfidence scoresHF ↗arXiv ↗
30

DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning

Yujie Wei, Xinyu Liu, Shiwei Zhang +12 authors

DreamVideo-Omni is a unified framework for video synthesis that enables precise multi-subject identity control and multi-granularity motion manipulation through a two-stage training approach combining condition-aware 3D positional embeddings, hierarchical motion injection, and latent identity reward learning.

32diffusion modelsvideo synthesisHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号