TensorX
返回文献探索

Paper · arXiv 2308.00566

Predicting masked tokens in stochastic locations improves masked image modeling

Amir Bar, Florian Bordes, Assaf Shocher, Mahmoud Assran, Pascal Vincent, Nicolas Ballas, Trevor Darrell, Amir Globerson, Yann LeCun

17 upvotesJuly 31, 2023arXiv 预印本
AI 摘要

FlexPredict enhances self-supervised learning in image processing by incorporating stochastic location uncertainty into masked image modeling, leading to improved performance on tasks like linear probing and video segmentation.

self-supervised learningpretext tasksmasked language modelingmasked image modelingFlexPredictstochastic modellocation uncertaintystochastic masked token positionsdownstream performanceImageNet linear probingsemi-supervised video segmentationViT-BViT-L

Abstract

Self-supervised learning is a promising paradigm in deep learning that enables learning from unlabeled data by constructing pretext tasks that require learning useful representations. In natural language processing, the dominant pretext task has been masked language modeling (MLM), while in computer vision there exists an equivalent called Masked Image Modeling (MIM). However, MIM is challenging because it requires predicting semantic content in accurate locations. E.g, given an incomplete picture of a dog, we can guess that there is a tail, but we cannot determine its exact location. In this work, we propose FlexPredict, a stochastic model that addresses this challenge by incorporating location uncertainty into the model. Specifically, we condition the model on stochastic masked token positions to guide the model toward learning features that are more robust to location uncertainties. Our approach improves downstream performance on a range of tasks, e.g, compared to MIM baselines, FlexPredict boosts ImageNet linear probing by 1.6% with ViT-B and by 2.5% for semi-supervised video segmentation using ViT-L.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号