5 papers
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
Marek Šuppa, Andrej Ridzik, Daniel Hládek +2
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- near…
SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
David Å tevaÅák, Marek Å uppa
Keyphrase extraction for morphologically rich, low-resource languages remains understudied, largely due to the scarcity of suitable evaluation datasets. We address this gap for Slo…
Enhancing BERT Fine-Tuning for Sentiment Analysis in Lower-Resourced Languages
Jozef KubÃk, Marek Å uppa, Martin TakáÄ
Limited data for low-resource languages typically yield weaker language models (LMs). Since pre-training is compute-intensive, it is more pragmatic to target improvements during fi…
skLEP: A Slovak General Language Understanding Benchmark
Marek Šuppa, Andrej Ridzik, Daniel Hládek +5
In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP…
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam +42
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage. While mu…