3 papers
cs.CL2026
Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment -- Is English Enough?
Adnan Al Ali, Kathy Hämmerl, Jindřich Libovický +1
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to E…
cs.CL2026
CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios
Michal Tichý, Jindřich Libovický
We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language identification: cousin languages…
cs.CL2025
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
Jindřich Libovický, Jindřich Helcl, Andrei Manea +1
We introduce CUS-QA, a benchmark for evaluation of open-ended regional question answering that encompasses both textual and visual modalities. We also provide strong baselines usin…