TensorX
返回文献探索

Paper · arXiv 2404.14396

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Yuying Ge, Sijie Zhao, Jinguo Zhu, Yixiao Ge, Kun Yi, Lin Song, Chen Li, Xiaohan Ding, Ying Shan

19 upvotesApril 22, 2024arXiv 预印本
AI 摘要

SEED-X is a versatile multimodal foundation model that enhances vision-language understanding and generation by comprehending diverse image sizes and enabling multi-granularity image generation.

multimodal foundation modelvision-language understandinggenerationSEED-LLaMASEED-Xmulti-granularity visual semantics

Abstract

The rapid evolution of multimodal foundation model has demonstrated significant progresses in vision-language understanding and generation, e.g., our previous work SEED-LLaMA. However, there remains a gap between its capability and the real-world applicability, primarily due to the model's limited capacity to effectively respond to various user instructions and interact with diverse visual data. In this work, we focus on bridging this gap through integrating two enhanced features: (1) comprehending images of arbitrary sizes and ratios, and (2) enabling multi-granularity image generation. We present a unified and versatile foundation model, namely, SEED-X, which is able to model multi-granularity visual semantics for comprehension and generation tasks. Besides the competitive results on public benchmarks, SEED-X demonstrates its effectiveness in handling real-world applications across various domains after instruction tuning. We hope that our work will inspire future research into what can be achieved by versatile multimodal foundation models in real-world applications. The models, codes, and datasets will be released in https://github.com/AILab-CVC/SEED-X.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation | TensorX