TensorX
返回文献探索

Paper · arXiv 2403.14611

Explorative Inbetweening of Time and Space

Haiwen Feng, Zheng Ding, Zhihao Xia, Simon Niklaus, Victoria Abrevaya, Michael J. Black, Xuaner Zhang

12 upvotesMarch 21, 2024arXiv 预印本
AI 摘要

Time Reversal Fusion enables video generation by fusing forward and backward denoising paths conditioned on start and end frames, outperforming existing methods in generating smooth, complex motions, and 3D-consistent views.

bounded generationimage-to-video modelTime Reversal Fusiondenoising pathsinbetweeningseamless video loopingevaluation dataset

Abstract

We introduce bounded generation as a generalized task to control video generation to synthesize arbitrary camera and subject motion based only on a given start and end frame. Our objective is to fully leverage the inherent generalization capability of an image-to-video model without additional training or fine-tuning of the original model. This is achieved through the proposed new sampling strategy, which we call Time Reversal Fusion, that fuses the temporally forward and backward denoising paths conditioned on the start and end frame, respectively. The fused path results in a video that smoothly connects the two frames, generating inbetweening of faithful subject motion, novel views of static scenes, and seamless video looping when the two bounding frames are identical. We curate a diverse evaluation dataset of image pairs and compare against the closest existing methods. We find that Time Reversal Fusion outperforms related work on all subtasks, exhibiting the ability to generate complex motions and 3D-consistent views guided by bounded frames. See project page at https://time-reversal.github.io.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Explorative Inbetweening of Time and Space | TensorX