S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
Jiajun Shi, Siyuan Tao, Yuhao Wu +18 authors
S³Gym evaluates whether large language model agents can self-test, self-judge, and self-improve through interaction experience across text-based games, revealing that effective self-improvement depends on task structure and experience representation.