From the 1 of 4 linked papers with an AI index.
4 papers
RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar
Marek Šuppa, Viktória Ondrejová, Lucia Ganajová +2
The paper presents RAGthoven, a multi-stage large language model pipeline that uses retrieval‑augmented generation and computational humor theories to produce constrained humor in…
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
Marek Šuppa, Andrej Ridzik, Daniel Hládek +2
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- near…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
skLEP: A Slovak General Language Understanding Benchmark
Marek Šuppa, Andrej Ridzik, Daniel Hládek +5
In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP…