activity
20192026
most citedHomograph Disambiguation Through Selective Diacritic Restoration

20 citations · 37 across the 9 of their papers we have counts for

collaborators

10 papers

cs.CL2026

Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs

Yara Alakeel, Chatrine Qwaider, Hanan Aldarmaki +1

This work investigates how effectively large language models (LLMs) and their tokenization schemes represent and generate Arabic root-pattern morphology, probing whether they captu…

cs.CL2026

Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models

Sawsan Alqahtani, Mir Tafseer Nayeem, Md Tahmid Rahman Laskar +2

Tokenization underlies every large language model, yet it remains an under-theorized and inconsistently designed component. Common subword approaches such as Byte Pair Encoding (BP…

cs.CL2025★ 1 cited

Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation

Mir Tafseer Nayeem, Sawsan Alqahtani, Md Tahmid Rahman Laskar +2

Tokenization is a crucial but under-evaluated step in large language models (LLMs). The standard metric, fertility (the average number of tokens per word), captures compression eff…

cs.CL2024★ 2 cited

A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations

Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari +10

Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough…

cs.CL2023

Automatic Restoration of Diacritics for Speech Data Sets

Sara Shatnawi, Sawsan Alqahtani, Hanan Aldarmaki

Automatic text-based diacritic restoration models generally have high diacritic error rates when applied to speech transcripts as a result of domain and style shifts in spoken lang…

cs.CL2022★ 1 cited

Injecting Domain Knowledge in Language Models for Task-Oriented Dialogue Systems

Denis Emelin, Daniele Bonadiman, Sawsan Alqahtani +2

Pre-trained language models (PLM) have advanced the state-of-the-art across NLP applications, but lack domain-specific knowledge that does not naturally occur in pre-training data.…