TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 6 – Apr 12, 2026

50 篇论文 · 按点赞排序

31

GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers

Shufan Jiang, Chios Chen, Zhiyang Chen

Large language models struggle with autonomous bug discovery in complex runtime environments, as demonstrated by a new game development benchmark that reveals limited effectiveness of current approaches despite sophisticated multi-agent systems and interactive agents.

48large language modelsautonomous bug discoveryHF ↗arXiv ↗
33

Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning

Qisheng Su, Shiting Huang, Zhen Fang +3 authors

Researchers introduce PTE (Prefill Token Equivalents), a hardware-aware metric for measuring efficiency in Tool-Integrated Reasoning scenarios, which better correlates with actual inference latency than traditional token counts by accounting for KV-Cache inefficiencies and long tool responses.

46Tool-Integrated ReasoningLLMsHF ↗arXiv ↗
35

Can LLMs Learn to Reason Robustly under Noisy Supervision?

Shenzhi Yang, Guangcheng Zhu, Bowen Song +7 authors

Reinforcement Learning with Verifiable Rewards faces challenges with noisy labels, but a proposed method called Online Label Refinement addresses this by progressively correcting labels based on policy improvement and consistency checks, demonstrating improved robustness across various mathematical reasoning benchmarks.

42Reinforcement Learning with Verifiable Rewardsnoisy labelsHF ↗arXiv ↗
39

FileGram: Grounding Agent Personalization in File-System Behavioral Traces

Shuai Liu, Shulin Tian, Kairui Hu +6 authors

FileGram is a framework for personalized AI agents that uses file-system behavioral traces to enhance memory systems and agent personalization, featuring a data engine, diagnostic benchmark, and memory architecture built from atomic actions and content changes.

40file-system behavioral tracespersona-driven data engineHF ↗arXiv ↗
42

MARS: Enabling Autoregressive Models Multi-Token Generation

Ziqi Jin, Lei Wang, Ziwei Luo +1 authors

MARS is a fine-tuning method that enables autoregressive language models to predict multiple tokens per forward pass without architectural changes, maintaining accuracy while improving throughput and supporting dynamic speed adjustment.

38autoregressive language modelstoken predictionHF ↗arXiv ↗
43

LightThinker++: From Reasoning Compression to Memory Management

Yuqi Zhu, Jintian Zhang, Zhenjie Wan +7 authors

LightThinker and LightThinker++ enable efficient large language model reasoning through dynamic compression and adaptive memory management, significantly reducing computational overhead while maintaining performance in complex tasks.

38large language modelsintermediate thoughtsHF ↗arXiv ↗
46

Watch Before You Answer: Learning from Visually Grounded Post-Training

Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang +8 authors

Vision-language models face challenges in video understanding due to text-based biases in benchmarks and datasets, which are addressed through VidGround, a method that uses only visually grounded questions for post-training to improve performance.

36vision-language modelsvideo understandingHF ↗arXiv ↗
48

SkillX: Automatically Constructing Skill Knowledge Bases for Agents

Chenxi Wang, Zhuoyun Yu, Xin Xie +8 authors

SkillX is an automated framework that creates reusable skill libraries for LLM agents through hierarchical skill design, iterative refinement, and exploratory expansion to improve generalization and efficiency across different environments.

35self-evolving paradigmsskill knowledge baseHF ↗arXiv ↗
49

Vero: An Open RL Recipe for General Visual Reasoning

Gabriel Sarch, Linrong Cai, Qunzhong Wang +3 authors

Vero is an open vision-language model family that achieves state-of-the-art visual reasoning performance through scaled reinforcement learning data across diverse tasks, demonstrating that broad data coverage drives strong RL scaling rather than isolated task-specific patterns.

35vision-language modelsreinforcement learningHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号