4 papers
Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment -- Is English Enough?
Adnan Al Ali, Kathy Hämmerl, Kathy Hämmerl +3
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to E…
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
Lukáš Eigler, JindÅich Libovický, David Hurych
Validating evaluation metrics for NLG typically relies on expensive and time-consuming human annotations, which predominantly exist only for English datasets. We propose LLM as a M…
CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios
Michal Tichý, JindÅich Libovický
We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language identification: cousin languages…
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
JindÅich Libovický, JindÅich Helcl, Andrei Manea +1
We introduce CUS-QA, a benchmark for evaluation of open-ended regional question answering that encompasses both textual and visual modalities. We also provide strong baselines usin…