TensorX
返回文献探索

Paper · arXiv 2411.19103

VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models

Jeongho Ju, Daeyoung Kim, SunYoung Park, Youngjune Kim

22 upvotesNovember 28, 2024arXiv 预印本
AI 摘要

An open-source Korean-English vision-language model, VARCO-VISION, is introduced and demonstrates superior performance in bilingual image-text tasks compared to similar models, incorporating a step-by-step training strategy and supporting grounding, referring, and OCR.

vision-language modelVARCO-VISIONstep-by-step training strategylinguisticvisual informationbilingual image-text understandinggeneration abilitiesgroundingreferringOCRevaluation datasetsclosed-setopenset benchmarks

Abstract

In this paper, we introduce an open-source Korean-English vision-language model (VLM), VARCO-VISION. We incorporate a step-by-step training strategy that allows a model learn both linguistic and visual information while preserving the backbone model's knowledge. Our model demonstrates outstanding performance in diverse settings requiring bilingual image-text understanding and generation abilities compared to models of similar size. VARCO-VISION is also capable of grounding, referring, and OCR, expanding its usage and potential applications for real-world scenarios. In addition to the model, we release five Korean evaluation datasets, including four closed-set and one openset benchmarks. We anticipate that our milestone will broaden the opportunities for AI researchers aiming to train VLMs. VARCO-VISION is available at https://huggingface.co/NCSOFT/VARCO-VISION-14B.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models | TensorX