TensorX
返回文献探索

Paper · arXiv 2603.13398

Qianfan-OCR: A Unified End-to-End Model for Document Intelligence

Daxiang Dong, Mingming Zheng, Dong Xu, Chunhua Luo, Bairong Zhuang, Yuxuan Li, Ruoyun He, Haoran Wang, Wenyu Zhang, Wenbo Wang, Yicheng Wang, Xue Xiong, Ayong Zheng, Xiaoying Zuo, Ziwei Ou, Jingnan Gu, Quanhao Guo, Jianmin Wu, Dawei Yin, Dou Shen

155 upvotesMarch 11, 2026arXiv 预印本
AI 摘要

Qianfan-OCR is a 4B-parameter vision-language model that unifies document parsing, layout analysis, and understanding while maintaining strong performance across multiple OCR benchmarks through its Layout-as-Thought mechanism.

vision-language modelend-to-end modelMarkdown conversionlayout analysisdocument understandingthink tokensstructured layout representationsbounding boxeselement typesreading orderlayout groundingOCR benchmarksVLMskey information extraction

Abstract

We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conversion and supports diverse prompt-driven tasks including table extraction, chart understanding, document QA, and key information extraction. To address the loss of explicit layout analysis in end-to-end OCR, we propose Layout-as-Thought, an optional thinking phase triggered by special think tokens that generates structured layout representations -- bounding boxes, element types, and reading order -- before producing final outputs, recovering layout grounding capabilities while improving accuracy on complex layouts. Qianfan-OCR ranks first among end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8), achieves competitive results on OCRBench, CCOCR, DocVQA, and ChartQA against general VLMs of comparable scale, and attains the highest average score on public key information extraction benchmarks, surpassing Gemini-3.1-Pro, Seed-2.0, and Qwen3-VL-235B. The model is publicly accessible via the Baidu AI Cloud Qianfan platform.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence | TensorX