activity
20242026
collaborators

11 papers

cs.CL2026

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl +14

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resourc…

cs.CL2026

Multilingual Vision-Language Models, A Survey

Andrei-Alexandru Manea, Jindřich Libovický

This survey examines multilingual vision-language models that process text and images across languages. We review 33 models and 23 benchmarks, spanning encoder-only and generative…

cs.CL2026

Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography

Gianluca Vico, Jindřich Libovický

We present a crowdsourced dataset for Piedmontese, an endangered Romance language of northwestern Italy. The dataset comprises 145 Italian-Piedmontese parallel sentences derived fr…

cs.CL2026

Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors

Adnan Al Ali, Jindřich Helcl, Jindřich Libovický

LLM-based assistants have been widely popularised after the release of ChatGPT. Concerns have been raised about their misuse in academia, given the difficulty of distinguishing bet…

cs.CL2026

On the Credibility of Evaluating LLMs using Survey Questions

Jindřich Libovický

Recent studies evaluate the value orientation of large language models (LLMs) using adapted social surveys, typically by prompting models with survey questions and comparing their…

cs.CL2026

Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features

Abishek Stephen, Jindřich Libovický

We present a novel metric for the evaluation of the morphological plausibility of subword segmentation. Unlike the typically used morpheme boundary or retrieval F-score, which requ…