3 papers
cs.CL2024
EgyBERT: A Large Language Model Pretrained on Egyptian Dialect Corpora
Faisal Qarah
This study presents EgyBERT, an Arabic language model pretrained on 10.4 GB of Egyptian dialectal texts. We evaluated EgyBERT's performance by comparing it with five other multidia…
cs.CL2024
SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
Faisal Qarah
In this paper, we introduce SaudiBERT, a monodialect Arabic language model pretrained exclusively on Saudi dialectal text. To demonstrate the model's effectiveness, we compared Sau…
cs.CL2024
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
Faisal Qarah
Arabic poetry, with its rich linguistic features and profound cultural significance, presents a unique challenge to the Natural Language Processing (NLP) field. The complexity of i…