TensorX
返回文献探索

Paper · arXiv 2309.07906

Generative Image Dynamics

Zhengqi Li, Richard Tucker, Noah Snavely, Aleksander Holynski

54 upvotesSeptember 14, 2023arXiv 预印本
AI 摘要

A frequency-coordinated diffusion sampling process is used to predict long-term motion representations for still images, enabling dynamic video creation and interactive scene manipulation.

frequency-coordinated diffusion sampling processneural stochastic motion textureFourier domaindense motion trajectoriesimage-based rendering moduledynamic videosinteractive scene manipulation

Abstract

We present an approach to modeling an image-space prior on scene dynamics. Our prior is learned from a collection of motion trajectories extracted from real video sequences containing natural, oscillating motion such as trees, flowers, candles, and clothes blowing in the wind. Given a single image, our trained model uses a frequency-coordinated diffusion sampling process to predict a per-pixel long-term motion representation in the Fourier domain, which we call a neural stochastic motion texture. This representation can be converted into dense motion trajectories that span an entire video. Along with an image-based rendering module, these trajectories can be used for a number of downstream applications, such as turning still images into seamlessly looping dynamic videos, or allowing users to realistically interact with objects in real pictures.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Generative Image Dynamics | TensorX