TensorX
返回文献探索

Paper · arXiv 2309.11499

DreamLLM: Synergistic Multimodal Comprehension and Creation

Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, Xiangwen Kong, Xiangyu Zhang, Kaisheng Ma, Li Yi

60 upvotesSeptember 20, 2023arXiv 预印本
AI 摘要

DreamLLM, a framework for Multimodal Large Language Models, directly samples in the multimodal space to enhance comprehension and creation synergy, enabling free-form interleaved content generation.

generative modelingmultimodal spaceCLIPfeature extractorsinterleaved documentsmultimodal distributionszero-shot multimodal generalist

Abstract

This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamLLM operates on two fundamental principles. The first focuses on the generative modeling of both language and image posteriors by direct sampling in the raw multimodal space. This approach circumvents the limitations and information loss inherent to external feature extractors like CLIP, and a more thorough multimodal understanding is obtained. Second, DreamLLM fosters the generation of raw, interleaved documents, modeling both text and image contents, along with unstructured layouts. This allows DreamLLM to learn all conditional, marginal, and joint multimodal distributions effectively. As a result, DreamLLM is the first MLLM capable of generating free-form interleaved content. Comprehensive experiments highlight DreamLLM's superior performance as a zero-shot multimodal generalist, reaping from the enhanced learning synergy.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
DreamLLM: Synergistic Multimodal Comprehension and Creation | TensorX