TensorX
返回文献探索

Paper · arXiv 2607.09657

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen

59 upvotesJuly 10, 2026arXiv 预印本
AI 摘要

Visual pretraining on document images outperforms text-only pretraining for scalable language intelligence without extracting text.

visual pretrainingfoundation modelsunsupervised visual pretrainingvisual documentstext-only pretraininglanguage intelligence

Abstract

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Scalable Visual Pretraining for Language Intelligence | TensorX