TensorX
返回文献探索

Paper · arXiv 2409.00558

Compositional 3D-aware Video Generation with LLM Director

Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen, Jiang Bian

15 upvotesAugust 31, 2024arXiv 预印本
AI 摘要

A novel text-to-video generation method decomposes concepts into 3D representations using LLMs, composes them with multi-modal LLM guidance, and refines with 2D diffusion models to produce high-fidelity videos with controlled motion and flexibility.

generative modelslarge-scale internet datatext-to-video generationLarge Language ModelsLLM3D representation2D diffusion modelsScore Distillation Sampling

Abstract

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the generated video, such as the motion and appearance of specific characters and the movement of viewpoints. In this work, we propose a novel paradigm that generates each concept in 3D representation separately and then composes them with priors from Large Language Models (LLM) and 2D diffusion models. Specifically, given an input textual prompt, our scheme consists of three stages: 1) We leverage LLM as the director to first decompose the complex query into several sub-prompts that indicate individual concepts within the video~(e.g., scene, objects, motions), then we let LLM to invoke pre-trained expert models to obtain corresponding 3D representations of concepts. 2) To compose these representations, we prompt multi-modal LLM to produce coarse guidance on the scales and coordinates of trajectories for the objects. 3) To make the generated frames adhere to natural image distribution, we further leverage 2D diffusion priors and use Score Distillation Sampling to refine the composition. Extensive experiments demonstrate that our method can generate high-fidelity videos from text with diverse motion and flexible control over each concept. Project page: https://aka.ms/c3v.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Compositional 3D-aware Video Generation with LLM Director | TensorX