TensorX
返回文献探索

Paper · arXiv 2306.00637

Wuerstchen: Efficient Pretraining of Text-to-Image Models

Pablo Pernias, Dominic Rampas, Marc Aubreville

13 upvotesJune 1, 2023arXiv 预印本
AI 摘要

Wuerstchen employs latent diffusion strategies for efficient text-to-image synthesis with reduced computational requirements and improved inference speed.

latent diffusion strategiestext-to-image synthesiscomputational accessibilityreal-time applications

Abstract

We introduce Wuerstchen, a novel technique for text-to-image synthesis that unites competitive performance with unprecedented cost-effectiveness and ease of training on constrained hardware. Building on recent advancements in machine learning, our approach, which utilizes latent diffusion strategies at strong latent image compression rates, significantly reduces the computational burden, typically associated with state-of-the-art models, while preserving, if not enhancing, the quality of generated images. Wuerstchen achieves notable speed improvements at inference time, thereby rendering real-time applications more viable. One of the key advantages of our method lies in its modest training requirements of only 9,200 GPU hours, slashing the usual costs significantly without compromising the end performance. In a comparison against the state-of-the-art, we found the approach to yield strong competitiveness. This paper opens the door to a new line of research that prioritizes both performance and computational accessibility, hence democratizing the use of sophisticated AI technologies. Through Wuerstchen, we demonstrate a compelling stride forward in the realm of text-to-image synthesis, offering an innovative path to explore in future research.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号