TensorX
返回文献探索

Paper · arXiv 2309.15564

Jointly Training Large Autoregressive Multimodal Models

Emanuele Aiello, Lili Yu, Yixin Nie, Armen Aghajanyan, Barlas Oguz

8 upvotesSeptember 27, 2023arXiv 预印本
AI 摘要

The Joint Autoregressive Mixture (JAM) framework integrates text and image generation models for high-quality multimodal output, exceeding previous methods with a specialized instruction-tuning strategy.

Joint Autoregressive MixtureJAM frameworkmultimodal outputsinstruction-tuningtext generationimage generation

Abstract

In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modalities into a single, robust model capable of generating seamless multimodal outputs remains a significant challenge. To address this gap, we present the Joint Autoregressive Mixture (JAM) framework, a modular approach that systematically fuses existing text and image generation models. We also introduce a specialized, data-efficient instruction-tuning strategy, tailored for mixed-modal generation tasks. Our final instruct-tuned model demonstrates unparalleled performance in generating high-quality multimodal outputs and represents the first model explicitly designed for this purpose.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Jointly Training Large Autoregressive Multimodal Models | TensorX