TensorX
返回文献探索

Paper · arXiv 2305.10722

Discriminative Diffusion Models as Few-shot Vision and Language Learners

Xuehai He, Weixi Feng, Tsu-Jui Fu, Varun Jampani, Arjun Akula, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Eric Wang

4 upvotesMay 18, 2023arXiv 预印本
AI 摘要

A novel method, Discriminative Stable Diffusion (DSD), leverages pre-trained text-to-image diffusion models with attention-based prompt learning to achieve superior performance in few-shot image-text matching tasks.

diffusion modelstext-to-image generationpre-trained diffusion modelsDiscriminative Stable Diffusion (DSD)cross-attention scoreattention-based prompt learningimage-text matchingfew-shot image-text matching

Abstract

Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified in text prompts, can we leverage the powerful representations learned by pre-trained diffusion models for discriminative tasks such as image-text matching? To answer this question, we propose a novel approach, Discriminative Stable Diffusion (DSD), which turns pre-trained text-to-image diffusion models into few-shot discriminative learners. Our approach uses the cross-attention score of a Stable Diffusion model to capture the mutual influence between visual and textual information and fine-tune the model via attention-based prompt learning to perform image-text matching. By comparing DSD with state-of-the-art methods on several benchmark datasets, we demonstrate the potential of using pre-trained diffusion models for discriminative tasks with superior results on few-shot image-text matching.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Discriminative Diffusion Models as Few-shot Vision and Language Learners | TensorX