TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 27 – May 3, 2026

50 篇论文 · 按点赞排序

31

The Last Human-Written Paper: Agent-Native Research Artifacts

Jiachen Liu, Jiaxin Pei, Jintao Huang +34 authors

Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with a machine-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities.

26HF ↗arXiv ↗
33

Why Fine-Tuning Encourages Hallucinations and How to Fix It

Guy Kaplan, Zorik Gekhman, Zhen Zhu +5 authors

Supervised fine-tuning in large language models can cause factual hallucinations due to knowledge degradation, which can be reduced through self-distillation regularization and parameter freezing techniques.

26supervised fine-tuninghallucinationsHF ↗arXiv ↗
35

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

Kehong Gong, Zhengyu Wen, Dao Thien Phong +10 authors

A fully end-to-end framework for arbitrary-skeleton motion capture that jointly optimizes video-to-pose and pose-to-rotation prediction while addressing rotation ambiguity through reference pose-rotation pairs and skeleton-aware attention mechanisms.

24Video-to-Pose networkinverse-kinematicsHF ↗arXiv ↗
37

Sapiens2

Rawal Khirodkar, He Wen, Julieta Martinez +3 authors

Sapiens2 is a high-resolution transformer model family for human-centric vision that achieves superior performance through combined pretraining objectives, large-scale human image datasets, and architectural improvements enabling detailed dense prediction and semantic understanding.

23transformersmasked image reconstructionHF ↗arXiv ↗
40

Step-level Optimization for Efficient Computer-use Agents

Jinbiao Wei, Kangqi Ni, Yilun Zhao +2 authors

Computer-use agents often rely on expensive multimodal models for every interaction, but a more efficient approach uses lightweight policies with risk detection monitors to escalate to stronger models only when needed.

19computer-use agentsgraphical user interfacesHF ↗arXiv ↗
43

Co-Director: Agentic Generative Video Storytelling

Yale Song, Yiwen Song, Nick Losier +13 authors

Co-Director presents a hierarchical multi-agent framework that formulates video storytelling as a global optimization problem, using multi-armed bandits and multimodal self-refinement to maintain semantic coherence and outperform existing approaches.

17diffusion modelsagentic pipelinesHF ↗arXiv ↗
44

Building a Precise Video Language with Human-AI Oversight

Zhiqiu Lin, Chancharik Mitra, Siyuan Cen +13 authors

Video-language models are enhanced through structured visual specifications and human-AI oversight frameworks that improve captioning accuracy and enable detailed video generation control.

17video-language modelsvideo captioningHF ↗arXiv ↗
46

Step-Audio-R1.5 Technical Report

Yuxin Zhang, Xiangyu Tony Zhang, Daijiao Liu +16 authors

Audio language models trained with reinforcement learning from verified rewards suffer from reduced conversational quality, prompting a shift toward reinforcement learning from human feedback for improved immersive dialogue experiences.

16Chain-of-ThoughtReinforcement Learning with Verified RewardsHF ↗arXiv ↗
50

Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

Bobo Li, Rui Wu, Zibo Ji +5 authors

Large language model agents exhibit cognitive bias where self-reflection and mutual auditing lead to inconsistent error attributions, which are addressed through a dialectical reasoning framework that promotes perspective-invariant decision making.

16large language model agentsmulti-agent frameworksHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号