6 citations · 10 across the 5 of their papers we have counts for
5 papers
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers +4
The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance…
Text Characterization Toolkit
Daniel Simig, Tianlu Wang, Verna Dankers +4
In NLP, models are usually evaluated by reporting single-number performance scores on a number of readily available benchmarks, without much deeper analysis. Here, we argue that -…
Using Linguistic Typology to Enrich Multilingual Lexicons: the Case of Lexical Gaps in Kinship
Temuulen Khishigsuren, Gábor Bella, Khuyagbaatar Batsuren +6
This paper describes a method to enrich lexical resources with content relating to linguistic diversity, based on knowledge from the field of lexical typology. We capture the pheno…
Language Diversity: Visible to Humans, Exploitable by Machines
Gábor Bella, Erdenebileg Byambadorj, Yamini Chandrashekar +3
The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as w…
Towards Algorithmic Transparency: A Diversity Perspective
Fausto Giunchiglia, Jahna Otterbacher, Styliani Kleanthous +4
As the role of algorithmic systems and processes increases in society, so does the risk of bias, which can result in discrimination against individuals and social groups. Research…