TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 20 – Jul 26, 2026
本周最热313

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

Fan Jiang, Zhaoxu Sun, Mengchao Wang +38 authors

ABot-World-0 is a real-time action-conditioned video world model that uses progressive distillation, long-horizon alignment, and a co-designed streaming stack to enable efficient, long-horizon interactive world generation.

action-conditioned video world modelWorldExplorerVLM-based assessmentODE distillationHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Shuqi Lu, Chaofan Li, Kun Luo +21 authors

AREX is a recursively self-improving deep research agent that verifies answers constraint-wise, compresses verified evidence into a compact state, and uses targeted follow-up research to refine results over long horizons.

155recursively self-improving agentsconstraint-wise verificationHF ↗arXiv ↗
08

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Xinyu Geng, Xuanhua He, Sixiang Chen +7 authors

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools. DeepSearch-World contains 420K multi-hop QA tasks constructed from entity-level random walks and supports key agentic cognitive behaviors useful for self-evolving, including progress verification, grounded reflection, and failure recovery. DeepSearch-Evolve iteratively performs trajectory generation, filtering, data mixing, and fine-tuning to train stronger agents. Without distillation from more capable models, DeepSearch-World-9B achieves competitive performance compared with open-source agents, reaching 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, showing that verifiable environments enable scalable self-evolution for long-horizon web agents. We will release the environment, 420K training pool, validation set, model, and code to facilitate future research on self-improving deep search agents.

10

Generative World Renderer at the Speed of Play

Guixu Lin, Zheng-Hui Huang, Siqi Yang +3 authors

AlayaRenderer-Flash accelerates a generative world renderer to real-time speeds via few-step autoregressive streaming and distilled codecs while preserving structured scene dynamics.

84generative world rendererG-bufferHF ↗arXiv ↗
12

Loop the Loopies!

Zitian Gao, Yilong Chen, Yihao Xiao +4 authors

Loopie is a looped Mixture-of-Experts Transformer that outperforms larger vanilla models at equal compute and achieves gold-medal reasoning on 2025 IMO and IPhO.

79Looped TransformerMixture-of-ExpertsHF ↗arXiv ↗
21

xHC: Expanded Hyper-Connections

Xiangdong Zhang, Xiaohan Qin, Sunan Zou +10 authors

xHC enables large-scale residual-stream expansion in transformers via sparse updates and temporal feature augmentation, improving efficiency and downstream performance.

56Hyper-ConnectionsManifold-Constrained Hyper-ConnectionsHF ↗arXiv ↗
22

Cura 1T: Specialized Model for Agentic Healthcare

actAVA AI, Haolin Chen, Leon Qi +8 authors

Cura 1T is a healthcare-specialized LLM trained via recursive self-improvement to jointly handle consultation, clinical reasoning, diagnosis, and EHR tool use without degrading other capabilities.

53healthcare AI agentsclinical reasoningHF ↗arXiv ↗
24

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

Dingyun Zhang, Lixue Gong, Wei Liu

A unified model integrates video and image generation and editing by synthesizing video training data from image pairs via temporal flow warping, aligning modalities through mimic losses, and internalizing instruction-based region localization via sense-related tasks and latent and attention losses.

53pixel-pair temporal warped flow fieldmodality mimic generation lossHF ↗arXiv ↗
25

Visual Contrastive Self-Distillation

Yijun Liang, Yunjie Tian, Yijiang Li +4 authors

Visual Contrastive Self-Distillation improves on-policy self-distillation by contrasting original and content-erased image conditions to generate stronger training signals without external teachers or privileged data.

52on-policy self-distillationself-distillationHF ↗arXiv ↗
29

On-Policy Delta Distillation

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1 authors

On-policy delta distillation improves reasoning transfer by using the difference between a tuned teacher and its base model as a reward signal, yielding stronger performance with brief post-training.

39on-policy distillationdelta signalHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号