TensorX
返回文献探索

Paper · arXiv 2411.11045

StableV2V: Stablizing Shape Consistency in Video-to-Video Editing

Chang Liu, Rui Li, Kaidong Zhang, Yunwei Lan, Dong Liu

11 upvotesNovember 17, 2024arXiv 预印本
AI 摘要

StableV2V is a shape-consistent video editing method that decomposes the editing process into sequential steps to align motions with user prompts, outperforming existing methods in visual consistency and efficiency.

video editingshape-consistencyalignmentvisual consistencyinference efficiencyDAVIS-Edit

Abstract

Recent advancements of generative AI have significantly promoted content creation and editing, where prevailing studies further extend this exciting progress to video editing. In doing so, these studies mainly transfer the inherent motion patterns from the source videos to the edited ones, where results with inferior consistency to user prompts are often observed, due to the lack of particular alignments between the delivered motions and edited contents. To address this limitation, we present a shape-consistent video editing method, namely StableV2V, in this paper. Our method decomposes the entire editing pipeline into several sequential procedures, where it edits the first video frame, then establishes an alignment between the delivered motions and user prompts, and eventually propagates the edited contents to all other frames based on such alignment. Furthermore, we curate a testing benchmark, namely DAVIS-Edit, for a comprehensive evaluation of video editing, considering various types of prompts and difficulties. Experimental results and analyses illustrate the outperforming performance, visual consistency, and inference efficiency of our method compared to existing state-of-the-art studies.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
StableV2V: Stablizing Shape Consistency in Video-to-Video Editing | TensorX