2 papers
cs.CL2026
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
Muhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan +13
There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus…
cs.CL2023
DeltaScore: Fine-Grained Story Evaluation with Perturbations
Zhuohan Xie, Miao Li, Trevor Cohn +1
Numerous evaluation metrics have been developed for natural language generation tasks, but their effectiveness in evaluating stories is limited as they are not specifically tailore…