TensorX
返回文献探索

Paper · arXiv 2411.04989

SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation

Koichi Namekata, Sherwin Bahmani, Ziyi Wu, Yash Kant, Igor Gilitschenski, David B. Lindell

14 upvotesNovember 7, 2024arXiv 预印本
AI 摘要

SG-I2V is a self-guided framework for image-to-video generation that achieves zero-shot control and competitive visual quality and motion fidelity without fine-tuning or external knowledge.

image-to-video generationdiffusion modelzero-shot controlvisual qualitymotion fidelity

Abstract

Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error, e.g., involving re-generating videos with different random seeds. Recent techniques address this issue by fine-tuning a pre-trained model to follow conditioning signals, such as bounding boxes or point trajectories. Yet, this fine-tuning procedure can be computationally expensive, and it requires datasets with annotated object motion, which can be difficult to procure. In this work, we introduce SG-I2V, a framework for controllable image-to-video generation that is self-guidedx2013offering zero-shot control by relying solely on the knowledge present in a pre-trained image-to-video diffusion model without the need for fine-tuning or external knowledge. Our zero-shot method outperforms unsupervised baselines while being competitive with supervised models in terms of visual quality and motion fidelity.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation | TensorX