TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 3 – Mar 9, 2025
本周最热113

START: Self-taught Reasoner with Tools

Chengpeng Li, Mingfeng Xue, Zhenru Zhang +7 authors

START integrates external tools into large reasoning models to enhance capabilities, using techniques like Hint-infer and Hint Rejection Sampling Fine-Tuning, achieving high performance across various benchmarks.

Large reasoning modelsChain-of-thoughtCoTself-checkingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

Visual-RFT: Visual Reinforcement Fine-Tuning

Ziyu Liu, Zeyi Sun, Yuhang Zang +5 authors

Visual Reinforcement Fine-Tuning (Visual-RFT) enhances Large Vision-Language Models (LVLMs) through reinforcement learning with visual perception verifiable rewards, achieving competitive performance in various visual and reasoning tasks with limited data.

86Reinforcement Fine-Tuning (RFT)Visual Reinforcement Fine-Tuning (Visual-RFT)HF ↗arXiv ↗
08

EgoLife: Towards Egocentric Life Assistant

Jingkang Yang, Shuai Liu, Hongming Guo +19 authors

EgoLife introduces EgoButler, combining EgoGPT and EgoRAG, to provide AI-powered assistance through wearable glasses using a comprehensive egocentric dataset.

48AI-powered wearable glassesegocentric video captureHF ↗arXiv ↗
09

Chain of Draft: Thinking Faster by Writing Less

Silei Xu, Wenhao Xie, Lingxiao Zhao +1 authors

Chain of Draft (CoD) improves the efficiency of Large Language Models (LLMs) in reasoning tasks by generating concise intermediate thoughts, enhancing accuracy while reducing token usage, cost, and latency.

48Chain-of-Thought (CoT)Chain of Draft (CoD)HF ↗arXiv ↗
16

Multi-Turn Code Generation Through Single-Step Rewards

Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen +3 authors

We address the problem of code generation from multi-turn execution feedback. Existing methods either generate code without feedback or use complex, hierarchical reinforcement learning to optimize multi-turn rewards. We propose a simple yet scalable approach, muCode, that solves multi-turn code generation using only single-step rewards. Our key insight is that code generation is a one-step recoverable MDP, where the correct code can be recovered from any intermediate code state in a single turn. muCode iteratively trains both a generator to provide code solutions conditioned on multi-turn execution feedback and a verifier to score the newly generated code. Experimental evaluations show that our approach achieves significant improvements over the state-of-the-art baselines. We provide analysis of the design choices of the reward models and policy, and show the efficacy of muCode at utilizing the execution feedback. Our code is available at https://github.com/portal-cornell/muCode.

32MDP$\mu$CodeHF ↗arXiv ↗
30

Wikipedia in the Era of LLMs: Evolution and Risks

Siming Huang, Yuliang Xu, Mingmeng Geng +2 authors

The analysis of Large Language Models' impact on Wikipedia shows a modest influence on article content and NLP tasks like machine translation and RAG, highlighting potential future risks.

22Large Language Models (LLMs)retrieval-augmented generation (RAG)HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号