TensorX
返回文献探索

Paper · arXiv 2407.06135

ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Ethan Chern, Jiadi Su, Yan Ma, Pengfei Liu

22 upvotesJuly 8, 2024arXiv 预印本
AI 摘要

Anole, an autoregressive large multimodal model based on Meta AI's Chameleon, supports efficient and high-quality interleaved image-text generation using native integration and an innovative fine-tuning strategy.

open-sourcemultimodal modelsLMMslarge language modelsLLMsdiffusion modelsnative integrationautoregressiveparameter-efficient fine-tuning

Abstract

Previous open-source large multimodal models (LMMs) have faced several limitations: (1) they often lack native integration, requiring adapters to align visual representations with pre-trained large language models (LLMs); (2) many are restricted to single-modal generation; (3) while some support multimodal generation, they rely on separate diffusion models for visual modeling and generation. To mitigate these limitations, we present Anole, an open, autoregressive, native large multimodal model for interleaved image-text generation. We build Anole from Meta AI's Chameleon, adopting an innovative fine-tuning strategy that is both data-efficient and parameter-efficient. Anole demonstrates high-quality, coherent multimodal generation capabilities. We have open-sourced our model, training framework, and instruction tuning data.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号