4 papers
When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates
Junhong Liang, Noor Abo Mokh, Bashar Alhafni
Arabic and Hebrew, as closely related Semitic languages, share a substantial lexicon of true cognates, misleading false friends, and modern loanwords. This overlap poses a challeng…
ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation
Joseph Marvin Imperial, Junhong Liang, Belal Shoer +9
When a text is translated, does the translation retain the complexity of the original? We introduce ComplexityMT, a new challenge for assessing how text complexity and machine tran…
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
Muhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan +13
There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus…
RAIR: Retrieval-Augmented Iterative Refinement for Chinese Spelling Correction
Junhong Liang, Yu Zhou
Chinese Spelling Correction (CSC) aims to detect and correct erroneous tokens in sentences. Traditional CSC focuses on equal length correction and uses pretrained language models (…