11 papers
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl +14
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resourc…
Multilingual Vision-Language Models, A Survey
Andrei-Alexandru Manea, JindÅich Libovický
This survey examines multilingual vision-language models that process text and images across languages. We review 33 models and 23 benchmarks, spanning encoder-only and generative…
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
Gianluca Vico, JindÅich Libovický
We present a crowdsourced dataset for Piedmontese, an endangered Romance language of northwestern Italy. The dataset comprises 145 Italian-Piedmontese parallel sentences derived fr…
Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors
Adnan Al Ali, JindÅich Helcl, JindÅich Libovický
LLM-based assistants have been widely popularised after the release of ChatGPT. Concerns have been raised about their misuse in academia, given the difficulty of distinguishing bet…
On the Credibility of Evaluating LLMs using Survey Questions
JindÅich Libovický
Recent studies evaluate the value orientation of large language models (LLMs) using adapted social surveys, typically by prompting models with survey questions and comparing their…
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
Abishek Stephen, JindÅich Libovický
We present a novel metric for the evaluation of the morphological plausibility of subword segmentation. Unlike the typically used morpheme boundary or retrieval F-score, which requ…