TensorX
返回文献探索

Paper · arXiv 2401.01827

Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions

David Junhao Zhang, Dongxu Li, Hung Le, Mike Zheng Shou, Caiming Xiong, Doyen Sahoo

16 upvotesJanuary 3, 2024arXiv 预印本
AI 摘要

Moonshot, a new video generation model, uses multimodal conditioning with images and text, along with an optional pre-trained ControlNet for geometry, to improve video quality and temporal consistency.

video diffusion modelsmultimodal inputsmultimodal video blockspatialtemporal layerscross-attention layerpre-trained ControlNetimage animationvideo editing

Abstract

Most existing video diffusion models (VDMs) are limited to mere text conditions. Thereby, they are usually lacking in control over visual appearance and geometry structure of the generated videos. This work presents Moonshot, a new video generation model that conditions simultaneously on multimodal inputs of image and text. The model builts upon a core module, called multimodal video block (MVB), which consists of conventional spatialtemporal layers for representing video features, and a decoupled cross-attention layer to address image and text inputs for appearance conditioning. In addition, we carefully design the model architecture such that it can optionally integrate with pre-trained image ControlNet modules for geometry visual conditions, without needing of extra training overhead as opposed to prior methods. Experiments show that with versatile multimodal conditioning mechanisms, Moonshot demonstrates significant improvement on visual quality and temporal consistency compared to existing models. In addition, the model can be easily repurposed for a variety of generative applications, such as personalized video generation, image animation and video editing, unveiling its potential to serve as a fundamental architecture for controllable video generation. Models will be made public on https://github.com/salesforce/LAVIS.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号