activity
20212024
most citedEvaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

6 citations · 10 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20246 cited

Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers +4

The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance…

cs.CL2022

Text Characterization Toolkit

Daniel Simig, Tianlu Wang, Verna Dankers +4

In NLP, models are usually evaluated by reporting single-number performance scores on a number of readily available benchmarks, without much deeper analysis. Here, we argue that -…

cs.CL20222 cited

Using Linguistic Typology to Enrich Multilingual Lexicons: the Case of Lexical Gaps in Kinship

Temuulen Khishigsuren, Gábor Bella, Khuyagbaatar Batsuren +6

This paper describes a method to enrich lexical resources with content relating to linguistic diversity, based on knowledge from the field of lexical typology. We capture the pheno…

cs.CL2022

Language Diversity: Visible to Humans, Exploitable by Machines

Gábor Bella, Erdenebileg Byambadorj, Yamini Chandrashekar +3

The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as w…

cs.CY20212 cited

Towards Algorithmic Transparency: A Diversity Perspective

Fausto Giunchiglia, Jahna Otterbacher, Styliani Kleanthous +4

As the role of algorithmic systems and processes increases in society, so does the risk of bias, which can result in discrimination against individuals and social groups. Research…