TensorX
返回文献探索

Paper · arXiv 2401.05735

Object-Centric Diffusion for Efficient Video Editing

Kumara Kahatapitiya, Adil Karjauv, Davide Abati, Fatih Porikli, Yuki M. Asano, Amirhossein Habibian

10 upvotesJanuary 11, 2024arXiv 预印本
AI 摘要

Object-Centric Diffusion reduces memory and computational cost in diffusion-based video editing by focusing on salient regions and merging redundant tokens in the background, achieving up to a 10x latency reduction with comparable quality.

diffusion-based video editingtextual edit promptsdiffusion inversioncross-frame attentionObject-Centric DiffusionOCDObject-Centric SamplingObject-Centric 3D Token Merging

Abstract

Diffusion-based video editing have reached impressive quality and can transform either the global style, local structure, and attributes of given video inputs, following textual edit prompts. However, such solutions typically incur heavy memory and computational costs to generate temporally-coherent frames, either in the form of diffusion inversion and/or cross-frame attention. In this paper, we conduct an analysis of such inefficiencies, and suggest simple yet effective modifications that allow significant speed-ups whilst maintaining quality. Moreover, we introduce Object-Centric Diffusion, coined as OCD, to further reduce latency by allocating computations more towards foreground edited regions that are arguably more important for perceptual quality. We achieve this by two novel proposals: i) Object-Centric Sampling, decoupling the diffusion steps spent on salient regions or background, allocating most of the model capacity to the former, and ii) Object-Centric 3D Token Merging, which reduces cost of cross-frame attention by fusing redundant tokens in unimportant background regions. Both techniques are readily applicable to a given video editing model without retraining, and can drastically reduce its memory and computational cost. We evaluate our proposals on inversion-based and control-signal-based editing pipelines, and show a latency reduction up to 10x for a comparable synthesis quality.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Object-Centric Diffusion for Efficient Video Editing | TensorX