TensorX
返回文献探索

Paper · arXiv 2410.21169

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Qintong Zhang, Victor Shea-Jay Huang, Bin Wang, Junyuan Zhang, Zhengren Wang, Hao Liang, Shawn Wang, Matthieu Lin, Wentao Zhang, Conghui He

30 upvotesOctober 28, 2024arXiv 预印本
AI 摘要

The survey reviews document parsing methodologies, including modular pipeline systems and end-to-end models, focusing on layout detection, content extraction, and multi-modal data integration, while highlighting challenges and future research directions.

layout detectioncontent extractiontexttablesmathematical expressionsmulti-modal data integrationlarge vision-language models

Abstract

Document parsing is essential for converting unstructured and semi-structured documents-such as contracts, academic papers, and invoices-into structured, machine-readable data. Document parsing extract reliable structured data from unstructured inputs, providing huge convenience for numerous applications. Especially with recent achievements in Large Language Models, document parsing plays an indispensable role in both knowledge base construction and training data generation. This survey presents a comprehensive review of the current state of document parsing, covering key methodologies, from modular pipeline systems to end-to-end models driven by large vision-language models. Core components such as layout detection, content extraction (including text, tables, and mathematical expressions), and multi-modal data integration are examined in detail. Additionally, this paper discusses the challenges faced by modular document parsing systems and vision-language models in handling complex layouts, integrating multiple modules, and recognizing high-density text. It emphasizes the importance of developing larger and more diverse datasets and outlines future research directions.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction | TensorX