TensorX
返回文献探索

Paper · arXiv 2403.03346

Enhancing Vision-Language Pre-training with Rich Supervisions

Yuan Gao, Kunyu Shi, Pengkai Zhu, Edouard Belval, Oren Nuriel, Srikar Appalaraju, Shabnam Ghadar, Vijay Mahadevan, Zhuowen Tu, Stefano Soatto

17 upvotesMarch 5, 2024arXiv 预印本
AI 摘要

Strongly Supervised pre-training with ScreenShots (S4) enhances Vision-Language Models by using web screenshots with tree-structured HTML elements and spatial localization, resulting in significant improvements in diverse downstream tasks.

Strongly Supervised pre-trainingScreenShotsVision-Language Modelsweb screenshotstree-structured hierarchyHTML elementsspatial localizationpre-training tasksimage-to-text modelTable DetectionWidget Captioning

Abstract

We propose Strongly Supervised pre-training with ScreenShots (S4) - a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. Using web screenshots unlocks a treasure trove of visual and textual cues that are not present in using image-text pairs. In S4, we leverage the inherent tree-structured hierarchy of HTML elements and the spatial localization to carefully design 10 pre-training tasks with large scale annotated data. These tasks resemble downstream tasks across different domains and the annotations are cheap to obtain. We demonstrate that, compared to current screenshot pre-training objectives, our innovative pre-training method significantly enhances performance of image-to-text model in nine varied and popular downstream tasks - up to 76.1% improvements on Table Detection, and at least 1% on Widget Captioning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Enhancing Vision-Language Pre-training with Rich Supervisions | TensorX