TensorX
返回文献探索

Paper · arXiv 2412.06578

MoViE: Mobile Diffusion for Video Editing

Adil Karjauv, Noor Fathima, Ioannis Lelekas, Fatih Porikli, Amir Ghodrati, Amirhossein Habibian

18 upvotesDecember 9, 2024arXiv 预印本
AI 摘要

Optimizations including a lightweight autoencoder, classifier-free guidance distillation, and an adversarial distillation scheme enable high-speed, high-quality video editing on mobile devices.

diffusion-based video editinglightweight autoencoderclassifier-free guidance distillationadversarial distillationmobile devices

Abstract

Recent progress in diffusion-based video editing has shown remarkable potential for practical applications. However, these methods remain prohibitively expensive and challenging to deploy on mobile devices. In this study, we introduce a series of optimizations that render mobile video editing feasible. Building upon the existing image editing model, we first optimize its architecture and incorporate a lightweight autoencoder. Subsequently, we extend classifier-free guidance distillation to multiple modalities, resulting in a threefold on-device speedup. Finally, we reduce the number of sampling steps to one by introducing a novel adversarial distillation scheme which preserves the controllability of the editing process. Collectively, these optimizations enable video editing at 12 frames per second on mobile devices, while maintaining high quality. Our results are available at https://qualcomm-ai-research.github.io/mobile-video-editing/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号