1 citations · 1 across the 6 of their papers we have counts for
8 papers · 1 filter
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
Chester Palen-Michel, Maxwell Pickering, Maya Kruse +2
We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotat…
Evaluating Morphological Compositional Generalization in Large Language Models
Mete Ismayilzada, Defne Circi, Jonne Sälevä +6
Large language models (LLMs) have demonstrated significant progress in various natural language generation and understanding tasks. However, their linguistic generalization capabil…
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata
Jonne Sälevä, Constantine Lignos
We introduce ParaNames, a massively multilingual parallel name resource consisting of 140 million names spanning over 400 languages. Names are provided for 16.8 million entities, a…
What changes when you randomly choose BPE merge operations? Not much
Jonne Sälevä, Constantine Lignos
We introduce three simple randomized variants of byte pair encoding (BPE) and explore whether randomizing the selection of merge operations substantially affects a downstream machi…
Toward More Meaningful Resources for Lower-resourced Languages
Constantine Lignos, Nolan Holley, Chester Palen-Michel +1
In this position paper, we describe our perspective on how meaningful resources for lower-resourced languages should be developed in connection with the speakers of those languages…
ParaNames: A Massively Multilingual Entity Name Corpus
Jonne Sälevä, Constantine Lignos
We introduce ParaNames, a multilingual parallel name resource consisting of 118 million names spanning across 400 languages. Names are provided for 13.6 million entities which are…