TensorX
返回文献探索

Paper · arXiv 2410.23287

ReferEverything: Towards Segmenting Everything We Can Speak of in Videos

Anurag Bagchi, Zhipeng Bao, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

19 upvotesOctober 30, 2024arXiv 预印本
AI 摘要

REM, a video segmentation framework using visual-language models, excels in segmenting and tracking unseen and dynamic concepts by leveraging Internet-scale pre-training.

visual-language representationsvideo diffusion modelsInternet-scale datasetsfine-tuningReferral Object Segmentationregion similaritysegmentingtrackingrare objectsunseen objectsdynamic conceptswaves crashingReferral Video Process SegmentationRef-DAVIS

Abstract

We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method capitalizes on visual-language representations learned by video diffusion models on Internet-scale datasets. A key insight of our approach is preserving as much of the generative model's original representation as possible, while fine-tuning it on narrow-domain Referral Object Segmentation datasets. As a result, our framework can accurately segment and track rare and unseen objects, despite being trained on object masks from a limited set of categories. Additionally, it can generalize to non-object dynamic concepts, such as waves crashing in the ocean, as demonstrated in our newly introduced benchmark for Referral Video Process Segmentation (Ref-VPS). Our experiments show that REM performs on par with state-of-the-art approaches on in-domain datasets, like Ref-DAVIS, while outperforming them by up to twelve points in terms of region similarity on out-of-domain data, leveraging the power of Internet-scale pre-training.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos | TensorX