TensorX
返回文献探索

Paper · arXiv 2407.20581

Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings

Gili Goldin, Shuly Wintner

24 upvotesJuly 30, 2024arXiv 预印本
AI 摘要

Knesset-DictaBERT, a Hebrew language model fine-tuned on Israeli parliamentary data, shows enhanced performance in understanding parliamentary language compared to the original DictaBERT.

Knesset-DictaBERTDictaBERTParliamentary languageMLM taskPerplexityAccuracy

Abstract

We present Knesset-DictaBERT, a large Hebrew language model fine-tuned on the Knesset Corpus, which comprises Israeli parliamentary proceedings. The model is based on the DictaBERT architecture and demonstrates significant improvements in understanding parliamentary language according to the MLM task. We provide a detailed evaluation of the model's performance, showing improvements in perplexity and accuracy over the baseline DictaBERT model.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings | TensorX