TensorX
返回文献探索

Paper · arXiv 2312.03511

Kandinsky 3.0 Technical Report

Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev, Anastasia Maltseva, Said Azizov, Igor Pavlov, Julia Agafonova, Andrey Kuznetsov, Denis Dimitrov

45 upvotesDecember 6, 2023arXiv 预印本
AI 摘要

Kandinsky 3.0, a large-scale text-to-image model based on latent diffusion, improves quality and realism through a larger architecture and advanced text understanding.

latent diffusionU-Nettext encoderdiffusion mappingtext-to-imagemodel architecturetext understandingspecific domains

Abstract

We present Kandinsky 3.0, a large-scale text-to-image generation model based on latent diffusion, continuing the series of text-to-image Kandinsky models and reflecting our progress to achieve higher quality and realism of image generation. Compared to previous versions of Kandinsky 2.x, Kandinsky 3.0 leverages a two times larger U-Net backbone, a ten times larger text encoder and removes diffusion mapping. We describe the architecture of the model, the data collection procedure, the training technique, and the production system of user interaction. We focus on the key components that, as we have identified as a result of a large number of experiments, had the most significant impact on improving the quality of our model compared to the others. By our side-by-side comparisons, Kandinsky becomes better in text understanding and works better on specific domains. Project page: https://ai-forever.github.io/Kandinsky-3

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号