TensorX
返回文献探索

Paper · arXiv 2608.26582

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

Gyouk Chu, Myeongho Jeon, Eunho Yang

45 upvotesAugust 27, 2026arXiv 预印本
AI 摘要

J-Zero enables self-improving language models across verifiable and unverifiable domains through adversarial co-evolution of a task generator, solver, and judge using predefined preference pairs.

self-evolving language modelsChallenger-Solver-Judge co-evolutionadversarial interactionpreference pairsJ-Zero

Abstract

Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号