TensorX
返回文献探索

Paper · arXiv 2312.02928

LivePhoto: Real Image Animation with Text-guided Motion Control

Xi Chen, Zhiheng Liu, Mengting Chen, Yutong Feng, Yu Liu, Yujun Shen, Hengshuang Zhao

18 upvotesDecember 5, 2023arXiv 预印本
AI 摘要

A text-to-video system, LivePhoto, uses a text-to-image generator enhanced with a motion module to accurately translate textual motion descriptions into videos.

text-to-video generationspatial contentstemporal motionstext-to-image generatorStable Diffusionmotion moduletemporal modelingtraining pipelinemotion intensity estimation moduletext re-weighting modulemotion-related textual instructionsvideo customization

Abstract

Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a challenge, this work presents a practical system, named LivePhoto, which allows users to animate an image of their interest with text descriptions. We first establish a strong baseline that helps a well-learned text-to-image generator (i.e., Stable Diffusion) take an image as a further input. We then equip the improved generator with a motion module for temporal modeling and propose a carefully designed training pipeline to better link texts and motions. In particular, considering the facts that (1) text can only describe motions roughly (e.g., regardless of the moving speed) and (2) text may include both content and motion descriptions, we introduce a motion intensity estimation module as well as a text re-weighting module to reduce the ambiguity of text-to-motion mapping. Empirical evidence suggests that our approach is capable of well decoding motion-related textual instructions into videos, such as actions, camera movements, or even conjuring new contents from thin air (e.g., pouring water into an empty glass). Interestingly, thanks to the proposed intensity learning mechanism, our system offers users an additional control signal (i.e., the motion intensity) besides text for video customization.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号