TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

686 篇论文 · 按点赞排序

331

SkillNet: Create, Evaluate, and Connect AI Skills

Yuan Liang, Ruobin Zhong, Haoming Xu +46 authors

SkillNet introduces an open infrastructure for systematically accumulating and transferring AI skills through a unified ontology, significantly improving agent performance across multiple domains.

95AI agentsskill consolidationHF ↗arXiv ↗
333

K-EXAONE Technical Report

Eunbi Choi, Kibong Choi, Seokhee Hong +62 authors

K-EXAONE is a multilingual language model with a Mixture-of-Experts architecture that achieves competitive performance on various benchmarks while supporting multiple languages and long-context windows.

95Mixture-of-Experts256K-token context windowHF ↗arXiv ↗
337

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

Haoyu Zhao, Xingyue Zhao, Siteng Huang +3 authors

A multi-modal 4D world model generates synchronized RGB, depth, and optical flow data from single RGB-D images and language instructions, enabling efficient robotic manipulation through unified diffusion processes and inverse dynamics policy learning.

95RGB-DFdiffusion modelsHF ↗arXiv ↗
340

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Xinyu Geng, Xuanhua He, Sixiang Chen +7 authors

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools. DeepSearch-World contains 420K multi-hop QA tasks constructed from entity-level random walks and supports key agentic cognitive behaviors useful for self-evolving, including progress verification, grounded reflection, and failure recovery. DeepSearch-Evolve iteratively performs trajectory generation, filtering, data mixing, and fine-tuning to train stronger agents. Without distillation from more capable models, DeepSearch-World-9B achieves competitive performance compared with open-source agents, reaching 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, showing that verifiable environments enable scalable self-evolution for long-horizon web agents. We will release the environment, 420K training pool, validation set, model, and code to facilitate future research on self-improving deep search agents.

347

Towards a Medical AI Scientist

Hongtao Wu, Boyun Zheng, Dingjie Song +5 authors

Medical AI Scientist represents the first autonomous research framework designed for clinical applications, enabling evidence-based hypothesis generation and manuscript drafting through clinician-engineer collaboration across three research modes.

92autonomous research frameworkclinical autonomous researchHF ↗arXiv ↗
354

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev +3 authors

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.2301, setting a new state of the art. Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.

92HF ↗arXiv ↗
356

SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

Ibragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov +1 authors

A large-scale dataset of software engineering tasks spanning multiple programming languages and repositories was created using an automated pipeline that generates executable environments and filters unreliable instances through LLM validation.

91reinforcement learningsoftware engineering agentsHF ↗arXiv ↗
12 / 23

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号