TensorX
返回文献探索

Paper · arXiv 2403.03883

SaulLM-7B: A pioneering Large Language Model for Law

Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Dominic Culver, Rui Melo, Caio Corro, Andre F. T. Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Morgado, Michael Desa

93 upvotesMarch 6, 2024arXiv 预印本
AI 摘要

SaulLM-7B, a large language model with 7 billion parameters, excels in legal text comprehension and generation using instructional fine-tuning on a legal corpus.

large language modellegal domainMistral 7B architectureEnglish legal corpusinstructional fine-tuninglegal datasets

Abstract

In this paper, we introduce SaulLM-7B, a large language model (LLM) tailored for the legal domain. With 7 billion parameters, SaulLM-7B is the first LLM designed explicitly for legal text comprehension and generation. Leveraging the Mistral 7B architecture as its foundation, SaulLM-7B is trained on an English legal corpus of over 30 billion tokens. SaulLM-7B exhibits state-of-the-art proficiency in understanding and processing legal documents. Additionally, we present a novel instructional fine-tuning method that leverages legal datasets to further enhance SaulLM-7B's performance in legal tasks. SaulLM-7B is released under the CC-BY-SA-4.0 License.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号