TensorX
返回文献探索

Paper · arXiv 2606.08671

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

Zhiwei Li, Yong Hu

48 upvotesJune 23, 2026arXiv 预印本
AI 摘要

SkillHone enables continuous evolution of agent skills by maintaining persistent decision histories and incorporating practice feedback for improved performance across research and tool-mediated analysis tasks.

agent skillsskill evolutiondecision historypractice feedbackcandidate skillsredacted reportingcross-session refinementdeep-research benchmarksGAIAWebWalkerQA-ENtool-mediated analysis

Abstract

Agent skills extend language-model agents with task-specific procedures, scripts, and references, but the tasks and environments they target continually change. Existing methods improve skills in bounded runs and retain only the final artifact, discarding the decision history that later agents need to interpret prior revisions, evaluations, and rejected alternatives. We introduce SkillHone, a harness for continual agent skill evolution grounded in persistent decision history. SkillHone pairs skill revisions with evaluation-side evidence that supplies practice feedback, recording structured histories of diagnoses, revisions, evidence, and outcomes. Role-separated subagents run candidate skills on practice probes with redacted reporting and propose revisions informed by prior decisions, enabling cross-session refinement without rediscovering past rationale. On deep-research benchmarks, SkillHone runs without a pre-integrated search stack and outperforms the commercially backed deep-research agent by 15.8 points on GAIA and 3.2 points on WebWalkerQA-EN, while also exceeding prior skill-evolution methods. We further deploy SkillHone on internal tool-mediated analysis scenarios, where it improves accuracy by an average of 18.8 points across seven settings.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History | TensorX