TensorX
返回文献探索

Paper · arXiv 2406.01300

pOps: Photo-Inspired Diffusion Operators

Elad Richardson, Yuval Alaluf, Ali Mahdavi-Amiri, Daniel Cohen-Or

17 upvotesJune 3, 2024arXiv 预印本
AI 摘要

pOps framework uses Diffusion Prior model to train semantic operators in CLIP image embedding space for text-guided image generation.

CLIP image embedding spaceIP-Adapterlinear operationspOpsDiffusion Prior modeldiffusion operatortextual CLIP lossphoto-inspired operators

Abstract

Text-guided image generation enables the creation of visual content from textual descriptions. However, certain visual concepts cannot be effectively conveyed through language alone. This has sparked a renewed interest in utilizing the CLIP image embedding space for more visually-oriented tasks through methods such as IP-Adapter. Interestingly, the CLIP image embedding space has been shown to be semantically meaningful, where linear operations within this space yield semantically meaningful results. Yet, the specific meaning of these operations can vary unpredictably across different images. To harness this potential, we introduce pOps, a framework that trains specific semantic operators directly on CLIP image embeddings. Each pOps operator is built upon a pretrained Diffusion Prior model. While the Diffusion Prior model was originally trained to map between text embeddings and image embeddings, we demonstrate that it can be tuned to accommodate new input conditions, resulting in a diffusion operator. Working directly over image embeddings not only improves our ability to learn semantic operations but also allows us to directly use a textual CLIP loss as an additional supervision when needed. We show that pOps can be used to learn a variety of photo-inspired operators with distinct semantic meanings, highlighting the semantic diversity and potential of our proposed approach.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
pOps: Photo-Inspired Diffusion Operators | TensorX