3 papers
cs.CL2023
People and Places of Historical Europe: Bootstrapping Annotation Pipeline and a New Corpus of Named Entities in Late Medieval Texts
Vít Novotný, Kristýna Luger, Michal Štefánik +2
Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. I…
cs.CL2023
Concept-aware Training Improves In-context Learning Ability of Language Models
Michal Štefánik, Marek Kadlčík
Many recent language models (LMs) of Transformers family exhibit so-called in-context learning (ICL) ability, manifested in the LMs' ability to modulate their function by a task de…
cs.CL2023
Resources and Few-shot Learners for In-context Learning in Slavic Languages
Michal Štefánik, Marek Kadlčík, Piotr Gramacki +1
Despite the rapid recent progress in creating accurate and compact in-context learners, most recent work focuses on in-context learning (ICL) for tasks in English. However, the abi…