TensorX
返回文献探索

Paper · arXiv 2504.10471

MIEB: Massive Image Embedding Benchmark

Chenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling, Xin Zhang, Márton Kardos, Roman Solomatin, Noura Al Moubayed, Kenneth Enevoldsen, Niklas Muennighoff

21 upvotesApril 14, 2025arXiv 预印本
AI 摘要

The Massive Image Embedding Benchmark (MIEB) evaluates image and image-text embedding models across various tasks, revealing hidden capabilities and performance correlations with multimodal large language models.

Massive Image Embedding BenchmarkMIEBimage embedding modeltext retrievalvision modelsimage-text embeddingmultilingualtask categoriesvisual representationinterleaved encodingsmultimodal large language models

Abstract

Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is equally good at retrieving relevant images given a piece of text. We introduce the Massive Image Embedding Benchmark (MIEB) to evaluate the performance of image and image-text embedding models across the broadest spectrum to date. MIEB spans 38 languages across 130 individual tasks, which we group into 8 high-level categories. We benchmark 50 models across our benchmark, finding that no single method dominates across all task categories. We reveal hidden capabilities in advanced vision models such as their accurate visual representation of texts, and their yet limited capabilities in interleaved encodings and matching images and texts in the presence of confounders. We also show that the performance of vision encoders on MIEB correlates highly with their performance when used in multimodal large language models. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号