collaborators

5 papers

cs.CL2026

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…

cs.CL2025

Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation

Jonne Sälevä, Duygu Ataman, Constantine Lignos

We introduce a set of resampling-based methods for quantifying uncertainty and statistical precision of evaluation metrics in multilingual and/or multitask NLP benchmarks. We show…

cs.CL2025

The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation

İbrahim Ethem Deveci, Duygu Ataman

The rapid rise of Large Language Models (LLMs) and Large Reasoning Models (LRMs) has been accompanied by an equally rapid increase of benchmarks used to assess them. However, due t…

cs.CL2025

TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages

Jafar Isbarov, Arofat Akhundjanova, Mammad Hajili +13

Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is essential for advancing the applicability of multilingual language models. However,…

cs.CL2025

Evaluating Morphological Compositional Generalization in Large Language Models

Mete Ismayilzada, Defne Circi, Jonne Sälevä +6

Large language models (LLMs) have demonstrated significant progress in various natural language generation and understanding tasks. However, their linguistic generalization capabil…