TensorX
返回文献探索

Paper · arXiv 2404.09204

TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Ya-Qi Yu, Minghui Liao, Jihao Wu, Yongxin Liao, Xiaoyu Zheng, Wei Zeng

11 upvotesApril 14, 2024arXiv 预印本
AI 摘要

TextHawk, a multimodal large language model designed for document-oriented tasks, outperforms existing models by using specialized components for efficient text and image perception.

ReSampling and ReArrangement (ReSA) moduleScalable Positional Embeddings (SPEs)Query Proposal Network (QPN)Multi-Level Cross-Attention (MLCA) mechanism

Abstract

Multimodal Large Language Models (MLLMs) have shown impressive results on various multimodal tasks. However, most existing MLLMs are not well suited for document-oriented tasks, which require fine-grained image perception and information compression. In this paper, we present TextHawk, a MLLM that is specifically designed for document-oriented tasks, while preserving the general capabilities of MLLMs. TextHawk is aimed to explore efficient fine-grained perception by designing four dedicated components. Firstly, a ReSampling and ReArrangement (ReSA) module is proposed to reduce the redundancy in the document texts and lower the computational cost of the MLLM. We explore encoding the positions of each local feature by presenting Scalable Positional Embeddings (SPEs), which can preserve the scalability of various image sizes. A Query Proposal Network (QPN) is then adopted to initialize the queries dynamically among different sub-images. To further enhance the fine-grained visual perceptual ability of the MLLM, we design a Multi-Level Cross-Attention (MLCA) mechanism that captures the hierarchical structure and semantic relations of document images. Furthermore, we create a new instruction-tuning dataset for document-oriented tasks by enriching the multimodal document data with Gemini Pro. We conduct extensive experiments on both general and document-oriented MLLM benchmarks, and show that TextHawk outperforms the state-of-the-art methods, demonstrating its effectiveness and superiority in fine-grained document perception and general abilities.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models | TensorX