TensorX
返回文献探索

Paper · arXiv 2412.14123

AnySat: An Earth Observation Model for Any Resolutions, Scales, and Modalities

Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu

11 upvotesDecember 18, 2024arXiv 预印本
AI 摘要

Anysat uses joint embedding predictive architecture and resolution-adaptive spatial encoders to train a unified multimodal model on diverse Earth observation datasets, achieving state-of-the-art performance across various environment monitoring tasks.

joint embedding predictive architectureJEPAresolution-adaptive spatial encodersmultimodal datasetsGeoPlexself-supervised mannerland cover mappingtree species identificationcrop type classificationchange detectionflood segmentation

Abstract

Geospatial models must adapt to the diversity of Earth observation data in terms of resolutions, scales, and modalities. However, existing approaches expect fixed input configurations, which limits their practical applicability. We propose AnySat, a multimodal model based on joint embedding predictive architecture (JEPA) and resolution-adaptive spatial encoders, allowing us to train a single model on highly heterogeneous data in a self-supervised manner. To demonstrate the advantages of this unified approach, we compile GeoPlex, a collection of 5 multimodal datasets with varying characteristics and 11 distinct sensors. We then train a single powerful model on these diverse datasets simultaneously. Once fine-tuned, we achieve better or near state-of-the-art results on the datasets of GeoPlex and 4 additional ones for 5 environment monitoring tasks: land cover mapping, tree species identification, crop type classification, change detection, and flood segmentation. The code and models are available at https://github.com/gastruc/AnySat.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
AnySat: An Earth Observation Model for Any Resolutions, Scales, and Modalities | TensorX