TensorX
返回文献探索

Paper · arXiv 2507.04569

Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts

Guokan Shang, Hadi Abdine, Ahmad Chamma, Amr Mohamed, Mohamed Anwar, Abdelaziz Bounhar, Omar El Herraoui, Preslav Nakov, Michalis Vazirgiannis, Eric Xing

24 upvotesJuly 6, 2025arXiv 预印本
AI 摘要

Nile-Chat models, using a Branch-Train-MiX strategy, outperform existing multilingual and Arabic LLMs on Egyptian dialect benchmarks, particularly in dual-script understanding and generation.

LLMsEgyptian dialectArabicLatin scriptsBranch-Train-MiXMoE modelmultilingualJaisALLaMEgyptian evaluation benchmarksQwen2.5-14B-Instruct

Abstract

We introduce Nile-Chat-4B, 3x4B-A6B, and 12B, a collection of LLMs for Egyptian dialect, uniquely designed to understand and generate texts written in both Arabic and Latin scripts. Specifically, with Nile-Chat-3x4B-A6B, we introduce a novel language adaptation approach by leveraging the Branch-Train-MiX strategy to merge script-specialized experts, into a single MoE model. Our Nile-Chat models significantly outperform leading multilingual and Arabic LLMs, such as LLaMa, Jais, and ALLaM, on our newly introduced Egyptian evaluation benchmarks, which span both understanding and generative tasks. Notably, our 12B model yields a 14.4% performance gain over Qwen2.5-14B-Instruct on Latin-script benchmarks. All our resources are publicly available. We believe this work presents a comprehensive methodology for adapting LLMs to dual-script languages, addressing an often overlooked aspect in modern LLM development.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号