GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
Shufan Jiang, Chios Chen, Zhiyang Chen
Large language models struggle with autonomous bug discovery in complex runtime environments, as demonstrated by a new game development benchmark that reveals limited effectiveness of current approaches despite sophisticated multi-agent systems and interactive agents.