TensorX
返回文献探索

Paper · arXiv 2505.04364

Benchmarking LLMs' Swarm intelligence

Kai Ruan, Mowen Huang, Ji-Rong Wen, Hao Sun

20 upvotesMay 7, 2025arXiv 预印本
AI 摘要

New benchmark evaluates LLMs in decentralized coordination tasks under limited information, highlighting challenges and potential for future systems.

Large Language ModelsMulti-Agent SystemsSwarmBenchdecentralized coordinationlocal perceptionlocal communicationswarm intelligencecoordination tasks2D grid environmentk x k viewcoordination effectivenessemergent group dynamicszero-shot evaluationrobust planningstrategy formationuncertaintyEmbodied MAS

Abstract

Large Language Models (LLMs) show potential for complex reasoning, yet their capacity for emergent coordination in Multi-Agent Systems (MAS) when operating under strict constraints-such as limited local perception and communication, characteristic of natural swarms-remains largely unexplored, particularly concerning the nuances of swarm intelligence. Existing benchmarks often do not fully capture the unique challenges of decentralized coordination that arise when agents operate with incomplete spatio-temporal information. To bridge this gap, we introduce SwarmBench, a novel benchmark designed to systematically evaluate the swarm intelligence capabilities of LLMs acting as decentralized agents. SwarmBench features five foundational MAS coordination tasks within a configurable 2D grid environment, forcing agents to rely primarily on local sensory input (k x k view) and local communication. We propose metrics for coordination effectiveness and analyze emergent group dynamics. Evaluating several leading LLMs in a zero-shot setting, we find significant performance variations across tasks, highlighting the difficulties posed by local information constraints. While some coordination emerges, results indicate limitations in robust planning and strategy formation under uncertainty in these decentralized scenarios. Assessing LLMs under swarm-like conditions is crucial for realizing their potential in future decentralized systems. We release SwarmBench as an open, extensible toolkit-built upon a customizable and scalable physical system with defined mechanical properties. It provides environments, prompts, evaluation scripts, and the comprehensive experimental datasets generated, aiming to foster reproducible research into LLM-based MAS coordination and the theoretical underpinnings of Embodied MAS. Our code repository is available at https://github.com/x66ccff/swarmbench.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Benchmarking LLMs' Swarm intelligence | TensorX