1 citations · 1 across the 5 of their papers we have counts for
5 papers
Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models
Reem I. Masoud, Chen Feng, Shunta Asano +3
The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning datasets used for cultural adapt…
BALSAM: A Platform for Benchmarking Arabic Large Language Models
Rawan Al-Matham, Kareem Darwish, Raghad Al-Rasheed +40
The impressive advancement of Large Language Models (LLMs) in English has not been matched across all languages. In particular, LLM performance in Arabic lags behind, due to data s…
Leveraging Corpus Metadata to Detect Template-based Translation: An Exploratory Case Study of the Egyptian Arabic Wikipedia Edition
Saied Alshahrani, Hesham Haroon, Ali Elfilali +2
Wikipedia articles (content pages) are commonly used corpora in Natural Language Processing (NLP) research, especially in low-resource languages other than English. Yet, a few rese…
Arabic Synonym BERT-based Adversarial Examples for Text Classification
Norah Alshahrani, Saied Alshahrani, Esma Wali +1
Text classification systems have been proven vulnerable to adversarial text examples, modified versions of the original text examples that are often unnoticed by human eyes, yet ca…
CIDAR: Culturally Relevant Instruction Dataset For Arabic
Zaid Alyafeai, Khalid Almubarak, Ahmed Ashraf +9
Instruction tuning has emerged as a prominent methodology for teaching Large Language Models (LLMs) to follow instructions. However, current instruction datasets predominantly cate…