TensorX
返回文献探索

Paper · arXiv 2403.09029

Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset

Hugo Laurençon, Léo Tronchon, Victor Sanh

57 upvotesMarch 14, 2024arXiv 预印本
AI 摘要

A synthetic dataset of HTML code and screenshots is introduced to enhance the ability of vision-language models to convert screenshots into HTML code.

vision-language modelsVLMsWebSightsynthetic datasetHTML codesscreenshotsfine-tune

Abstract

Using vision-language models (VLMs) in web development presents a promising strategy to increase efficiency and unblock no-code solutions: by providing a screenshot or a sketch of a UI, a VLM could generate the code to reproduce it, for instance in a language like HTML. Despite the advancements in VLMs for various tasks, the specific challenge of converting a screenshot into a corresponding HTML has been minimally explored. We posit that this is mainly due to the absence of a suitable, high-quality dataset. This work introduces WebSight, a synthetic dataset consisting of 2 million pairs of HTML codes and their corresponding screenshots. We fine-tune a foundational VLM on our dataset and show proficiency in converting webpage screenshots to functional HTML code. To accelerate the research in this area, we open-source WebSight.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset | TensorX