4 papers
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report
Mariia Fedorova, Nikolay Arefyev, Maja Buljan +4
Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to i…
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
Jindřich Libovický, Jindřich Helcl, Andrei Manea +1
We introduce CUS-QA, a benchmark for evaluation of open-ended regional question answering that encompasses both textual and visual modalities. We also provide strong baselines usin…
Atyaephyra at SemEval-2025 Task 4: Low-Rank Negative Preference Optimization
Jan Bronec, Jindřich Helcl
We present a submission to the SemEval 2025 shared task on unlearning sensitive content from LLMs. Our approach employs negative preference optimization using low-rank adaptation.…
Teaching LLMs at Charles University: Assignments and Activities
Jindřich Helcl, Zdeněk Kasner, Ondřej Dušek +4
This paper presents teaching materials, particularly assignments and ideas for classroom activities, from a new course on large language models (LLMs) taught at Charles University.…