TensorX
返回文献探索

Paper · arXiv 2503.15485

TULIP: Towards Unified Language-Image Pretraining

Zineng Tang, Long Lian, Seun Eisape, XuDong Wang, Roei Herzig, Adam Yala, Alane Suhr, Trevor Darrell, David M. Chan

49 upvotesMarch 19, 2025arXiv 预印本
AI 摘要

TULIP enhances image-text contrastive models by incorporating generative data augmentation and contrastive learning to improve fine-grained visual features and zero-shot performance.

CLIPSigLIPvision-centric taskshigh-fidelity image understandingcountingdepth estimationfine-grained object recognitionlanguage alignmentgenerative data augmentationimage-image contrastive learningtext-text contrastive learningimage/text reconstruction regularizationImageNet-1KRxRx1zero-shot performancefew-shot classificationvision-language modelsMMVP

Abstract

Despite the recent success of image-text contrastive models like CLIP and SigLIP, these models often struggle with vision-centric tasks that demand high-fidelity image understanding, such as counting, depth estimation, and fine-grained object recognition. These models, by performing language alignment, tend to prioritize high-level semantics over visual understanding, weakening their image understanding. On the other hand, vision-focused models are great at processing visual information but struggle to understand language, limiting their flexibility for language-driven tasks. In this work, we introduce TULIP, an open-source, drop-in replacement for existing CLIP-like models. Our method leverages generative data augmentation, enhanced image-image and text-text contrastive learning, and image/text reconstruction regularization to learn fine-grained visual features while preserving global semantic alignment. Our approach, scaling to over 1B parameters, outperforms existing state-of-the-art (SOTA) models across multiple benchmarks, establishing a new SOTA zero-shot performance on ImageNet-1K, delivering up to a 2times enhancement over SigLIP on RxRx1 in linear probing for few-shot classification, and improving vision-language models, achieving over 3times higher scores than SigLIP on MMVP. Our code/checkpoints are available at https://tulip-berkeley.github.io

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号