TensorX
返回文献探索

Paper · arXiv 2606.26740

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

Xinyu Wang, Chongbo Zhao, Fangneng Zhan, Yue Ma

83 upvotesJune 25, 2026arXiv 预印本
AI 摘要

A novel streaming video editing framework enables causal, frame-by-frame editing with stable long-horizon preservation and real-time responsiveness through a three-stage distillation pipeline and AR-oriented mask cache.

streaming video editingcausal editingframe-by-frame editingcontent preservationreal-time responsivenessthree-stage distillation pipelinebidirectional foundation modelunidirectional streaming editorlong-horizon editsAR-oriented mask cacheinference speedinteractive applicationsaugmented reality

Abstract

Streaming video editing has made rapid progress, yet practical deployment is still limited by two core issues: maintaining stable backgrounds and non-edited regions over time, and achieving the low latency required for real-time interactive scenarios. Meanwhile, recent streaming video generation methods are mostly developed for synthesis and cannot be directly applied to editing due to the strict preservation requirement and region-specific control. In this work, we present a novel streaming video editing framework that performs causal, frame-by-frame editing with strong content preservation and real-time responsiveness. Our key design is a three-stage distillation pipeline that progressively transfers editing capability from a powerful bidirectional foundation model to an efficient unidirectional streaming editor, enabling stable long-horizon edits without sacrificing visual fidelity. To further support real-time deployment, we introduce an AR-oriented mask cache that reuses region-related computation across frames, substantially reducing redundant processing and accelerating inference. Finally, we establish a dedicated benchmark for streaming video editing. Extensive evaluations demonstrate that our method achieves state-of-the-art visual quality among streaming baselines while drastically boosting inference speed to 12.66 FPS, making it suitable for interactive and augmented reality applications.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing | TensorX