TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热779

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

Björn Engdahl, Adrian Kosowski, Jan Chorowski +6 authors

A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1.

in-context learningrecurrent latent reasoninghigh-dimensional latent spaceARC-AGI-1HF ↗arXiv ↗

50 篇论文 · 按点赞排序

10

Recursive Synthesis for Long-Horizon Terminal Tasks

Zhongzhi Li, Yucheng Shi, Zongxia Li +8 authors

Recursive verified synthesis generates scalable long-horizon terminal-agent training data, substantially improving model performance on terminal benchmarks through supervised fine-tuning and reinforcement learning.

252recursive verified synthesisterminal-agent tasksHF ↗arXiv ↗
11

On-Policy Self-Distillation without Any Supervision

Yijiang Li, Bingyang Wang, Yijun Liang +3 authors

Unsupervised on-policy self-distillation improves large language models by using internal consistency and majority-vote pseudo-solutions to correct confident errors without external supervision.

219on-policy self-distillationself-consistencyHF ↗arXiv ↗
12

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex Team, B. An, B. Li +68 authors

Apodex 1.1 improves sustained, verifiable progress on complex real-world tasks by scaling executable environments and training agents to coordinate long-horizon work with state maintenance and recovery.

207agentic coordination scalingenvironment scalingHF ↗arXiv ↗
15

Beyond Pixels: From Video Priors to 4D Worlds

Zihao Liu, Xiaolong Shen, Zhenglin Zhou +2 authors

Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining.

1894D generationvariational autoencoderHF ↗arXiv ↗
19

Demystifying Agent Skills: Why They Work-Until They Don't

Zhiyuan Jiang, Fangrui Huang, Hanwen Xing +6 authors

Skills enhance LLM agents primarily by stabilizing execution through procedural anchoring rather than injecting missing knowledge, though retrieval bottlenecks and brittle assumptions limit their effectiveness.

170LLM agentsskillsHF ↗arXiv ↗
20

Self-Supervised Visual On-Policy Distillation

Yijiang Li, Yijun Liang, Yunjie Tian +6 authors

Self-supervised visual on-policy distillation improves small vision-language models by distilling from original images into strongly augmented student views without privileged annotations or larger teachers.

170visual on-policy distillationteacher-student asymmetryHF ↗arXiv ↗
24

DAPD: Dual-Anchored Policy Distillation

Jianyu Wu, Yizhou Wang, Encheng Su +2 authors

Dual-Anchored Policy Distillation resolves privilege illusion in on-policy self-distillation by aligning matched-information paths and reducing reliance on privileged guidance, improving performance across model scales.

153on-policy self-distillationprivilege illusionHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号