TensorX
返回文献探索

Paper · arXiv 2403.00504

Learning and Leveraging World Models in Visual Representation Learning

Quentin Garrido, Mahmoud Assran, Nicolas Ballas, Adrien Bardes, Laurent Najman, Yann LeCun

33 upvotesMarch 1, 2024arXiv 预印本
AI 摘要

Image World Models extend Joint-Embedding Predictive Architecture to predict global photometric transformations, offering adaptable and high-performance representations through fine-tuning.

Joint-Embedding Predictive ArchitectureImage World Modelsmasked image modelingglobal photometric transformationslatent spaceconditioningprediction difficultycapacityfine-tuningcontrastive methodsequivariant representations

Abstract

Joint-Embedding Predictive Architecture (JEPA) has emerged as a promising self-supervised approach that learns by leveraging a world model. While previously limited to predicting missing parts of an input, we explore how to generalize the JEPA prediction task to a broader set of corruptions. We introduce Image World Models, an approach that goes beyond masked image modeling and learns to predict the effect of global photometric transformations in latent space. We study the recipe of learning performant IWMs and show that it relies on three key aspects: conditioning, prediction difficulty, and capacity. Additionally, we show that the predictive world model learned by IWM can be adapted through finetuning to solve diverse tasks; a fine-tuned IWM world model matches or surpasses the performance of previous self-supervised methods. Finally, we show that learning with an IWM allows one to control the abstraction level of the learned representations, learning invariant representations such as contrastive methods, or equivariant representations such as masked image modelling.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号