TensorX
返回文献探索

Paper · arXiv 2310.12274

An Image is Worth Multiple Words: Learning Object Level Concepts using Multi-Concept Prompt Learning

Chen Jin, Ryutaro Tanno, Amrutha Saseendran, Tom Diethe, Philip Teare

13 upvotesOctober 18, 2023arXiv 预印本
AI 摘要

A novel Multi-Concept Prompt Learning framework (MCPL) learns multiple embeddings from a single sentence-image pair to represent object-level concepts in images, using regularisation techniques to enhance accuracy.

prompt learningTextural Inversionimage styleimage appearanceMulti-Concept Prompt LearningMCPLAttention MaskingAttnMaskPrompts Contrastive LossPromptCLBind adjectiveBind adj.word-concept correlationimage generationimage editingattention visualisationsemantic disentangled conceptsdatasetevaluation protocol

Abstract

Textural Inversion, a prompt learning method, learns a singular embedding for a new "word" to represent image style and appearance, allowing it to be integrated into natural language sentences to generate novel synthesised images. However, identifying and integrating multiple object-level concepts within one scene poses significant challenges even when embeddings for individual concepts are attainable. This is further confirmed by our empirical tests. To address this challenge, we introduce a framework for Multi-Concept Prompt Learning (MCPL), where multiple new "words" are simultaneously learned from a single sentence-image pair. To enhance the accuracy of word-concept correlation, we propose three regularisation techniques: Attention Masking (AttnMask) to concentrate learning on relevant areas; Prompts Contrastive Loss (PromptCL) to separate the embeddings of different concepts; and Bind adjective (Bind adj.) to associate new "words" with known words. We evaluate via image generation, editing, and attention visualisation with diverse images. Extensive quantitative comparisons demonstrate that our method can learn more semantically disentangled concepts with enhanced word-concept correlation. Additionally, we introduce a novel dataset and evaluation protocol tailored for this new task of learning object-level concepts.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
An Image is Worth Multiple Words: Learning Object Level Concepts using Multi-Concept Prompt Learning | TensorX