TensorX
返回文献探索

Paper · arXiv 2411.14793

Style-Friendly SNR Sampler for Style-Driven Generation

Jooyoung Choi, Chaehun Shin, Yeongtak Oh, Heeseung Kim, Sungroh Yoon

40 upvotesNovember 22, 2024arXiv 预印本
AI 摘要

The Style-friendly SNR sampler modifies the noise level distribution during fine-tuning to improve style alignment in diffusion models, enabling better capture of unique artistic styles.

diffusion modelssignal-to-noise ratio (SNR)fine-tuningnoise level distributionstylistic featuresstyle templatespersonalized content creation

Abstract

Recent large-scale diffusion models generate high-quality images but struggle to learn new, personalized artistic styles, which limits the creation of unique style templates. Fine-tuning with reference images is the most promising approach, but it often blindly utilizes objectives and noise level distributions used for pre-training, leading to suboptimal style alignment. We propose the Style-friendly SNR sampler, which aggressively shifts the signal-to-noise ratio (SNR) distribution toward higher noise levels during fine-tuning to focus on noise levels where stylistic features emerge. This enables models to better capture unique styles and generate images with higher style alignment. Our method allows diffusion models to learn and share new "style templates", enhancing personalized content creation. We demonstrate the ability to generate styles such as personal watercolor paintings, minimal flat cartoons, 3D renderings, multi-panel images, and memes with text, thereby broadening the scope of style-driven generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号