TensorX
返回文献探索

Paper · arXiv 2408.16766

CSGO: Content-Style Composition in Text-to-Image Generation

Peng Xing, Haofan Wang, Yanpeng Sun, Qixun Wang, Xu Bai, Hao Ai, Renyuan Huang, Zechao Li

18 upvotesSeptember 4, 2024arXiv 预印本
AI 摘要

A large-scale style transfer dataset and end-to-end training model are presented that effectively decouple content and style features for enhanced image generation control.

diffusion modelimage style transferimage inversionstyle transfer datasetcontent-style-stylized image tripletsend-to-end trainingfeature injection

Abstract

The diffusion model has shown exceptional capabilities in controlled image generation, which has further fueled interest in image style transfer. Existing works mainly focus on training free-based methods (e.g., image inversion) due to the scarcity of specific data. In this study, we present a data construction pipeline for content-style-stylized image triplets that generates and automatically cleanses stylized data triplets. Based on this pipeline, we construct a dataset IMAGStyle, the first large-scale style transfer dataset containing 210k image triplets, available for the community to explore and research. Equipped with IMAGStyle, we propose CSGO, a style transfer model based on end-to-end training, which explicitly decouples content and style features employing independent feature injection. The unified CSGO implements image-driven style transfer, text-driven stylized synthesis, and text editing-driven stylized synthesis. Extensive experiments demonstrate the effectiveness of our approach in enhancing style control capabilities in image generation. Additional visualization and access to the source code can be located on the project page: https://csgo-gen.github.io/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
CSGO: Content-Style Composition in Text-to-Image Generation | TensorX