TensorX
返回文献探索

Paper · arXiv 2306.05178

SyncDiffusion: Coherent Montage via Synchronized Joint Diffusions

Yuseung Lee, Kunho Kim, Hyunjin Kim, Minhyuk Sung

7 upvotesJune 8, 2023arXiv 预印本
AI 摘要

SyncDiffusion synchronizes multiple diffusions via perceptual similarity loss to generate coherent panoramic images without visible seams and maintains fidelity and compatibility with input prompts.

pretrained image diffusion modelspanoramasnaive stitchingseamless montage generationjoint diffusionslatent featuresperceptual similarity lossgradient descentdenoised imagesGIQACLIP score

Abstract

The remarkable capabilities of pretrained image diffusion models have been utilized not only for generating fixed-size images but also for creating panoramas. However, naive stitching of multiple images often results in visible seams. Recent techniques have attempted to address this issue by performing joint diffusions in multiple windows and averaging latent features in overlapping regions. However, these approaches, which focus on seamless montage generation, often yield incoherent outputs by blending different scenes within a single image. To overcome this limitation, we propose SyncDiffusion, a plug-and-play module that synchronizes multiple diffusions through gradient descent from a perceptual similarity loss. Specifically, we compute the gradient of the perceptual loss using the predicted denoised images at each denoising step, providing meaningful guidance for achieving coherent montages. Our experimental results demonstrate that our method produces significantly more coherent outputs compared to previous methods (66.35% vs. 33.65% in our user study) while still maintaining fidelity (as assessed by GIQA) and compatibility with the input prompt (as measured by CLIP score).

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号