TensorX
返回文献探索

Paper · arXiv 2410.18538

SMITE: Segment Me In TimE

Amirhossein Alimohammadi, Sauradip Nag, Saeid Asgari Taghanaki, Andrea Tagliasacchi, Ghassan Hamarneh, Ali Mahdavi Amiri

15 upvotesOctober 24, 2024arXiv 预印本
AI 摘要

A pre-trained text-to-image diffusion model with tracking mechanism outperforms existing methods in object segmentation across video frames with arbitrary granularity.

diffusion modeltracking mechanismobject segmentationvideo framesarbitrary granularity

Abstract

Segmenting an object in a video presents significant challenges. Each pixel must be accurately labelled, and these labels must remain consistent across frames. The difficulty increases when the segmentation is with arbitrary granularity, meaning the number of segments can vary arbitrarily, and masks are defined based on only one or a few sample images. In this paper, we address this issue by employing a pre-trained text to image diffusion model supplemented with an additional tracking mechanism. We demonstrate that our approach can effectively manage various segmentation scenarios and outperforms state-of-the-art alternatives.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号