HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
Petr Anokhin, Roman Khalikov, Stefan Rebrikov +3 authors
HeroBench evaluates long-horizon planning and structured reasoning in complex virtual worlds, revealing performance disparities in state-of-the-art LLMs.