TensorX
返回文献探索

Paper · arXiv 2411.09703

MagicQuill: An Intelligent Interactive Image Editing System

Zichen Liu, Yue Yu, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Wen Wang, Zhiheng Liu, Qifeng Chen, Yujun Shen

80 upvotesNovember 14, 2024arXiv 预印本
AI 摘要

MagicQuill integrates an interface with a multimodal large language model and diffusion prior to enable efficient and precise real-time image editing through minimal user input.

multimodal large language modeldiffusion priortwo-branch plug-in module

Abstract

Image editing involves a variety of complex tasks and requires efficient and precise manipulation techniques. In this paper, we present MagicQuill, an integrated image editing system that enables swift actualization of creative ideas. Our system features a streamlined yet functionally robust interface, allowing for the articulation of editing operations (e.g., inserting elements, erasing objects, altering color) with minimal input. These interactions are monitored by a multimodal large language model (MLLM) to anticipate editing intentions in real time, bypassing the need for explicit prompt entry. Finally, we apply a powerful diffusion prior, enhanced by a carefully learned two-branch plug-in module, to process editing requests with precise control. Experimental results demonstrate the effectiveness of MagicQuill in achieving high-quality image edits. Please visit https://magic-quill.github.io to try out our system.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号