TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

32

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Ruiyan Gong, Yingnan Guo, Junjun Hu +43 authors

ABot-N1 improves visual language navigation by separating reasoning from control through a slow-fast architecture that uses explicit chain-of-thought reasoning and pixel goals to guide continuous waypoint generation, achieving strong results across diverse embodied tasks.

103Visual Language NavigationChain-of-Thought reasoningHF ↗arXiv ↗
35

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

Haoyu Zhao, Xingyue Zhao, Siteng Huang +3 authors

A multi-modal 4D world model generates synchronized RGB, depth, and optical flow data from single RGB-D images and language instructions, enabling efficient robotic manipulation through unified diffusion processes and inverse dynamics policy learning.

96RGB-DFdiffusion modelsHF ↗arXiv ↗
36

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Xinyu Geng, Xuanhua He, Sixiang Chen +7 authors

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools. DeepSearch-World contains 420K multi-hop QA tasks constructed from entity-level random walks and supports key agentic cognitive behaviors useful for self-evolving, including progress verification, grounded reflection, and failure recovery. DeepSearch-Evolve iteratively performs trajectory generation, filtering, data mixing, and fine-tuning to train stronger agents. Without distillation from more capable models, DeepSearch-World-9B achieves competitive performance compared with open-source agents, reaching 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, showing that verifiable environments enable scalable self-evolution for long-horizon web agents. We will release the environment, 420K training pool, validation set, model, and code to facilitate future research on self-improving deep search agents.

40

Video Generation Models are General-Purpose Vision Learners

Letian Wang, Chuhan Zhang, Rishabh Kabra +9 authors

GenCeption uses a video generative diffusion backbone to build a generalist vision model that achieves strong performance across diverse perception tasks with high data efficiency and emergent generalization.

90text-to-video generationgenerative diffusionHF ↗arXiv ↗
41

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Chen Tang, Yizhou Wang, Jianyu Wu +26 authors

SciReasoner is a multimodal scientific foundation model that enables interpretable structural reasoning across proteins, molecules, and crystals by discretizing structural elements into a unified vocabulary for enhanced prediction and scientific inference.

90multimodal scientific foundation modelstructural reasoningHF ↗arXiv ↗
48

Generative World Renderer at the Speed of Play

Guixu Lin, Zheng-Hui Huang, Siqi Yang +3 authors

AlayaRenderer-Flash accelerates a generative world renderer to real-time speeds via few-step autoregressive streaming and distilled codecs while preserving structured scene dynamics.

84generative world rendererG-bufferHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号