3 papers
cs.CL2026
Training data generation for context-dependent rubric-based short answer grading
Pavel Å indeláÅ, Dávid Slivka, Christopher Bouma +2
Every four years, the PISA test is administered by the OECD to test the knowledge of teenage students worldwide and allow for comparisons of educational systems. However, having to…
cs.CL2025
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
Pavel Å indeláÅ, OndÅej Bojar
ELOQUENT is a set of shared tasks that aims to create easily testable high-level criteria for evaluating generative language models. Sensemaking is one such shared task. In Sensema…
cs.CL2025
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
Petra BaranÄÃková, OndÅej Bojar
In this paper, we compare Czech-specific and multilingual sentence embedding models through intrinsic and extrinsic evaluation paradigms. For intrinsic evaluation, we employ Costra…