TensorX
返回文献探索

Paper · arXiv 2608.24804

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav, Sai Rajeswar, Patrice Bechard, Sridhar Nemala, Sagar Davasam

41 upvotesAugust 25, 2026arXiv 预印本
AI 摘要

StarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.

agent harnessesMCP-backed providerssubagent structureagent-loop configurationevolution poolstratifying tasksbaseline failure behaviorproposer-visible search tasksproposer-hidden selection tasksheld-out tasksITBench SREEnterpriseOps-Gym ITSMAutomationBench Financeharness evolutionGPTQwenfalse-positive diagnosesmodel-environment mismatch

Abstract

We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration. StarHarness constructs a compact evolution pool by stratifying tasks according to baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution improves full-benchmark performance by 20-35 percentage points over the default harness after 4-12 accepted changes per environment. These gains persist on tasks excluded from evolution and transfer without re-evolution across GPT and Qwen model families. Trace analysis links the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, with fewer false-positive diagnoses and shorter trajectories in several settings. StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号