TensorX
返回文献探索

Paper · arXiv 2408.09085

Segment Anything with Multiple Modalities

Aoran Xiao, Weihao Xuan, Heli Qi, Yun Xing, Naoto Yokoya, Shijian Lu

22 upvotesAugust 17, 2024arXiv 预印本
AI 摘要

MM-SAM extends SAM to support multi-modal and cross-modal processing, enhancing segmentation accuracy across various sensor modalities through unsupervised cross-modal transfer and weakly-supervised multi-modal fusion.

Segment Anything ModelSAMMM-SAMcross-modal transferweakly-supervised multi-modal fusionmulti-modal datasensor fusionmask-free trainingsingle-modal processingsensor modalities

Abstract

Robust and accurate segmentation of scenes has become one core functionality in various visual recognition and navigation tasks. This has inspired the recent development of Segment Anything Model (SAM), a foundation model for general mask segmentation. However, SAM is largely tailored for single-modal RGB images, limiting its applicability to multi-modal data captured with widely-adopted sensor suites, such as LiDAR plus RGB, depth plus RGB, thermal plus RGB, etc. We develop MM-SAM, an extension and expansion of SAM that supports cross-modal and multi-modal processing for robust and enhanced segmentation with different sensor suites. MM-SAM features two key designs, namely, unsupervised cross-modal transfer and weakly-supervised multi-modal fusion, enabling label-efficient and parameter-efficient adaptation toward various sensor modalities. It addresses three main challenges: 1) adaptation toward diverse non-RGB sensors for single-modal processing, 2) synergistic processing of multi-modal data via sensor fusion, and 3) mask-free training for different downstream tasks. Extensive experiments show that MM-SAM consistently outperforms SAM by large margins, demonstrating its effectiveness and robustness across various sensors and data modalities.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Segment Anything with Multiple Modalities | TensorX