8 papers
When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates
Junhong Liang, Noor Abo Mokh, Bashar Alhafni
Arabic and Hebrew, as closely related Semitic languages, share a substantial lexicon of true cognates, misleading false friends, and modern loanwords. This overlap poses a challeng…
Arabic Sentence Segmentation Across Genres and Punctuation Conditions
Mohammed Elkholy, Khalid N. Elmadani, Nizar Habash +1
Sentence segmentation in Arabic is challenging due to ambiguous and inconsistent punctuation, with many texts lacking reliable sentence boundary markers. Existing approaches rely h…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models
Mohamed Adel, Bashar Alhafni, Nizar Habash
Large language models (LLMs) perform strongly on many NLP tasks, but their ability to produce explicit linguistic structure remains unclear. We evaluate instruction-tuned LLMs on t…
Opportunities and Challenges of LLMs in Education: An NLP Perspective
Sowmya Vajjala, Bashar Alhafni, Stefano Bannò +2
Interest in the role of large language models (LLMs) in education is increasing, considering the new opportunities they offer for teaching, learning, and assessment. In this paper,…
BALSAM: A Platform for Benchmarking Arabic Large Language Models
Rawan Al-Matham, Kareem Darwish, Raghad Al-Rasheed +40
The impressive advancement of Large Language Models (LLMs) in English has not been matched across all languages. In particular, LLM performance in Arabic lags behind, due to data s…