TensorX
返回文献探索

Paper · arXiv 2412.07112

Maya: An Instruction Finetuned Multilingual Multimodal Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, S M Iftekhar Uddin, Shayekh Bin Islam, Roshan Santhosh, Snegha A, Drishti Sharma, Chen Liu, Isha Chaturvedi, Genta Indra Winata, Ashvanth. S, Snehanshu Mukherjee, Alham Fikri Aji

28 upvotesDecember 10, 2024arXiv 预印本
AI 摘要

Maya is an open-source multimodal multilingual model that addresses the limitations of Vision-Language Models in handling low-resource languages and cultural nuances by introducing a multilingual pretraining dataset and analyzing/removing toxicity.

Vision-Language ModelsVLMsmultilingual image-text pretraining datasetLLaVAtoxicity-free datasetmultilingual image-text modelcultural comprehensionlinguistic comprehension

Abstract

The rapid development of large Vision-Language Models (VLMs) has led to impressive results on academic benchmarks, primarily in widely spoken languages. However, significant gaps remain in the ability of current VLMs to handle low-resource languages and varied cultural contexts, largely due to a lack of high-quality, diverse, and safety-vetted data. Consequently, these models often struggle to understand low-resource languages and cultural nuances in a manner free from toxicity. To address these limitations, we introduce Maya, an open-source Multimodal Multilingual model. Our contributions are threefold: 1) a multilingual image-text pretraining dataset in eight languages, based on the LLaVA pretraining dataset; 2) a thorough analysis of toxicity within the LLaVA dataset, followed by the creation of a novel toxicity-free version across eight languages; and 3) a multilingual image-text model supporting these languages, enhancing cultural and linguistic comprehension in vision-language tasks. Code available at https://github.com/nahidalam/maya.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Maya: An Instruction Finetuned Multilingual Multimodal Model | TensorX