TensorX
返回文献探索

Paper · arXiv 2312.00738

SeaLLMs -- Large Language Models for Southeast Asia

Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, Hang Zhang, Lidong Bing

24 upvotesDecember 1, 2023arXiv 预印本
AI 摘要

SeaLLMs, a series of Llama-2 based models tailored for Southeast Asian languages, demonstrate superior performance across linguistic tasks and outperform ChatGPT-3.5 in non-Latin languages while being lightweight and cost-effective.

SeaLLMsLlama-2pre-trainingvocabularyinstruction tuningalignment tuninglinguistic tasksassistant-styleChatGPT-3.5ThaiKhmerLaoBurmese

Abstract

Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To address this imbalance, we introduce SeaLLMs, an innovative series of language models that specifically focuses on Southeast Asian (SEA) languages. SeaLLMs are built upon the Llama-2 model and further advanced through continued pre-training with an extended vocabulary, specialized instruction and alignment tuning to better capture the intricacies of regional languages. This allows them to respect and reflect local cultural norms, customs, stylistic preferences, and legal considerations. Our comprehensive evaluation demonstrates that SeaLLM-13b models exhibit superior performance across a wide spectrum of linguistic tasks and assistant-style instruction-following capabilities relative to comparable open-source models. Moreover, they outperform ChatGPT-3.5 in non-Latin languages, such as Thai, Khmer, Lao, and Burmese, by large margins while remaining lightweight and cost-effective to operate.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SeaLLMs -- Large Language Models for Southeast Asia | TensorX