TensorX
返回文献探索

Paper · arXiv 2312.14216

DreamDistribution: Prompt Distribution Learning for Text-to-Image Diffusion Models

Brian Nlong Zhao, Yuhang Xiao, Jiashu Xu, Xinyang Jiang, Yifan Yang, Dongsheng Li, Laurent Itti, Vibhav Vineet, Yunhao Ge

12 upvotesDecember 21, 2023arXiv 预印本
AI 摘要

A pretrained Text-to-Image diffusion model learns soft prompts to generate diverse images with specific attributes, adapting to text-to-3D and validated through quantitative and qualitative evaluations.

Text-to-Image (T2I) diffusion modelssoft promptslearned distributiontext-guided editingvariation controltext-to-3D

Abstract

The popularization of Text-to-Image (T2I) diffusion models enables the generation of high-quality images from text descriptions. However, generating diverse customized images with reference visual attributes remains challenging. This work focuses on personalizing T2I diffusion models at a more abstract concept or category level, adapting commonalities from a set of reference images while creating new instances with sufficient variations. We introduce a solution that allows a pretrained T2I diffusion model to learn a set of soft prompts, enabling the generation of novel images by sampling prompts from the learned distribution. These prompts offer text-guided editing capabilities and additional flexibility in controlling variation and mixing between multiple distributions. We also show the adaptability of the learned prompt distribution to other tasks, such as text-to-3D. Finally we demonstrate effectiveness of our approach through quantitative analysis including automatic evaluation and human assessment. Project website: https://briannlongzhao.github.io/DreamDistribution

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
DreamDistribution: Prompt Distribution Learning for Text-to-Image Diffusion Models | TensorX