5 papers
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
Chester Palen-Michel, Maxwell Pickering, Maya Kruse +2
We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotat…
Evaluating Morphological Compositional Generalization in Large Language Models
Mete Ismayilzada, Defne Circi, Jonne Sälevä +6
Large language models (LLMs) have demonstrated significant progress in various natural language generation and understanding tasks. However, their linguistic generalization capabil…
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata
Jonne Sälevä, Constantine Lignos
We introduce ParaNames, a massively multilingual parallel name resource consisting of 140 million names spanning over 400 languages. Names are provided for 16.8 million entities, a…
Toward More Meaningful Resources for Lower-resourced Languages
Constantine Lignos, Nolan Holley, Chester Palen-Michel +1
In this position paper, we describe our perspective on how meaningful resources for lower-resourced languages should be developed in connection with the speakers of those languages…
Mining Wikidata for Name Resources for African Languages
Jonne Sälevä, Constantine Lignos
This work supports further development of language technology for the languages of Africa by providing a Wikidata-derived resource of name lists corresponding to common entity type…