TensorX
返回文献探索

Paper · arXiv 2503.24379

Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation

Shengqiong Wu, Weicai Ye, Jiahao Wang, Quande Liu, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Shuicheng Yan, Hao Fei, Tat-Seng Chua

76 upvotesMarch 31, 2025arXiv 预印本
AI 摘要

Any2Caption framework improves video generation quality and controllability by decoupling condition interpretation and leveraging multimodal large language models.

multimodal large language modelsMLLMscaptioningtask decouplingvideo synthesisdense captionsstructured captionsvideo generatorsinstruction tuningAny2CapInslarge-scale datasetcontrollabilityvideo qualitytask evaluation

Abstract

To address the bottleneck of accurate user intent interpretation within the current video generation community, we present Any2Caption, a novel framework for controllable video generation under any condition. The key idea is to decouple various condition interpretation steps from the video synthesis step. By leveraging modern multimodal large language models (MLLMs), Any2Caption interprets diverse inputs--text, images, videos, and specialized cues such as region, motion, and camera poses--into dense, structured captions that offer backbone video generators with better guidance. We also introduce Any2CapIns, a large-scale dataset with 337K instances and 407K conditions for any-condition-to-caption instruction tuning. Comprehensive evaluations demonstrate significant improvements of our system in controllability and video quality across various aspects of existing video generation models. Project Page: https://sqwu.top/Any2Cap/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation | TensorX