TensorX
返回文献探索

Paper · arXiv 2409.08513

Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection

Haoxuan Wang, Qingdong He, Jinlong Peng, Hao Yang, Mingmin Chi, Yabiao Wang

13 upvotesSeptember 13, 2024arXiv 预印本
AI 摘要

Mamba-YOLO-World improves open-vocabulary detection by introducing a linear complexity MambaFusion-PAN neck architecture, enhancing performance on COCO and LVIS benchmarks.

open-vocabulary detectionOVDYOLO seriesYOLO-Worldneck feature fusionquadratic complexityguided receptive fieldsMambaFusion Path Aggregation NetworkMambaFusion-PANState Space ModelParallel-Guided Selective ScanSerial-Guided Selective Scanmulti-modal input sequencesmamba hidden statesselective scanning process

Abstract

Open-vocabulary detection (OVD) aims to detect objects beyond a predefined set of categories. As a pioneering model incorporating the YOLO series into OVD, YOLO-World is well-suited for scenarios prioritizing speed and efficiency.However, its performance is hindered by its neck feature fusion mechanism, which causes the quadratic complexity and the limited guided receptive fields.To address these limitations, we present Mamba-YOLO-World, a novel YOLO-based OVD model employing the proposed MambaFusion Path Aggregation Network (MambaFusion-PAN) as its neck architecture. Specifically, we introduce an innovative State Space Model-based feature fusion mechanism consisting of a Parallel-Guided Selective Scan algorithm and a Serial-Guided Selective Scan algorithm with linear complexity and globally guided receptive fields. It leverages multi-modal input sequences and mamba hidden states to guide the selective scanning process.Experiments demonstrate that our model outperforms the original YOLO-World on the COCO and LVIS benchmarks in both zero-shot and fine-tuning settings while maintaining comparable parameters and FLOPs. Additionally, it surpasses existing state-of-the-art OVD methods with fewer parameters and FLOPs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号