TensorX
返回文献探索

Paper · arXiv 2403.01487

InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding

Haogeng Liu, Quanzeng You, Xiaotian Han, Yiqi Wang, Bohan Zhai, Yongfei Liu, Yunzhe Tao, Huaibo Huang, Ran He, Hongxia Yang

15 upvotesMarch 3, 2024arXiv 预印本
AI 摘要

InfiMM-HD, a novel architecture with cross-attention and visual windows, enhances multimodal large language models for high-resolution image processing with efficient training and low computational overhead.

multimodal large language modelsMLLMsInfiMM-HDcross-attention modulevisual windowsfour-stage training pipeline

Abstract

Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intricate details within high-resolution images. Despite being indispensable for the development of robust MLLMs, this area remains underinvestigated. To tackle this challenge, our work introduces InfiMM-HD, a novel architecture specifically designed for processing images of different resolutions with low computational overhead. This innovation facilitates the enlargement of MLLMs to higher-resolution capabilities. InfiMM-HD incorporates a cross-attention module and visual windows to reduce computation costs. By integrating this architectural design with a four-stage training pipeline, our model attains improved visual perception efficiently and cost-effectively. Empirical study underscores the robustness and effectiveness of InfiMM-HD, opening new avenues for exploration in related areas. Codes and models can be found at https://huggingface.co/Infi-MM/infimm-hd

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号