3 papers
cs.CL2026
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
Marek Šuppa, Andrej Ridzik, Daniel Hládek +2
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- near…
cs.CL2025
o-MEGA: Optimized Methods for Explanation Generation and Analysis
ĽuboÅ¡ KriÅ¡, Jaroslav KopÄan, Qiwei Peng +3
The proliferation of transformer-based language models has revolutionized NLP domain while simultaneously introduced significant challenges regarding model transparency and trustwo…
cs.CL2025
skLEP: A Slovak General Language Understanding Benchmark
Marek Šuppa, Andrej Ridzik, Daniel Hládek +5
In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP…