TensorX
返回文献探索

Paper · arXiv 2406.07547

Zero-shot Image Editing with Reference Imitation

Xi Chen, Yutong Feng, Mengting Chen, Yiyang Wang, Shilong Zhang, Yu Liu, Yujun Shen, Hengshuang Zhao

31 upvotesJune 11, 2024arXiv 预印本
AI 摘要

MimicBrush, a generative training framework developed from a diffusion prior, enables imitative image editing by leveraging video frames and masked regions to capture semantic correspondence and improve editing quality.

generative training frameworkMimicBrushdiffusion priorvideo framesmasked regionssemantic correspondenceimitative editing

Abstract

Image editing serves as a practical yet challenging task considering the diverse demands from users, where one of the hardest parts is to precisely describe how the edited image should look like. In this work, we present a new form of editing, termed imitative editing, to help users exercise their creativity more conveniently. Concretely, to edit an image region of interest, users are free to directly draw inspiration from some in-the-wild references (e.g., some relative pictures come across online), without having to cope with the fit between the reference and the source. Such a design requires the system to automatically figure out what to expect from the reference to perform the editing. For this purpose, we propose a generative training framework, dubbed MimicBrush, which randomly selects two frames from a video clip, masks some regions of one frame, and learns to recover the masked regions using the information from the other frame. That way, our model, developed from a diffusion prior, is able to capture the semantic correspondence between separate images in a self-supervised manner. We experimentally show the effectiveness of our method under various test cases as well as its superiority over existing alternatives. We also construct a benchmark to facilitate further research.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Zero-shot Image Editing with Reference Imitation | TensorX