collaborators

7 papers

cs.CL2026

Last Translation Benchmark

Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…

cs.CL2026

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl +14

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resourc…

cs.CL2026

Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions

Antonia Karamolegkou, Nicolas Angleraud, Benoît Sagot +1

Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on l…

cs.CL2025

EvalCards: A Framework for Standardized Evaluation Reporting

Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou +11

Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Draw…

cs.CL2025

Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection

Ivan Vykopal, Antonia Karamolegkou, Jaroslav Kopčan +4

Multilingual Large Language Models (LLMs) offer powerful capabilities for cross-lingual fact-checking. However, these models often exhibit language bias, performing disproportionat…

cs.CL2025

Trick or Neat: Adversarial Ambiguity and Language Model Evaluation

Antonia Karamolegkou, Oliver Eberle, Phillip Rust +2

Detecting ambiguity is important for language understanding, including uncertainty estimation, humour detection, and processing garden path sentences. We assess language models' se…