most citedEvaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

6 citations · 17 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20246 cited

Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers +4

The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance…

cs.CL20241 cited

Advancing the Arabic WordNet: Elevating Content Quality

Abed Alhakim Freihat, Hadi Khalilia, Gábor Bella +1

High-quality WordNets are crucial for achieving high-quality results in NLP applications that rely on such resources. However, the wordnets of most languages suffer from serious is…

cs.CL20231 cited

Lexical Diversity in Kinship Across Languages and Dialects

Hadi Khalilia, Gábor Bella, Abed Alhakim Freihat +2

Languages are known to describe the world in diverse ways. Across lexicons, diversity is pervasive, appearing through phenomena such as lexical gaps and untranslatability. However,…

cs.CL20234 cited

Towards Bridging the Digital Language Divide

Gábor Bella, Paula Helm, Gertraud Koch +1

It is a well-known fact that current AI-based language technology -- language models, machine translation systems, multilingual dictionaries and corpora -- focuses on the world's 2…

cs.CL2023

Representing Interlingual Meaning in Lexical Databases

Fausto Giunchiglia, Gabor Bella, Nandu Chandran Nair +2

In today's multilingual lexical databases, the majority of the world's languages are under-represented. Beyond a mere issue of resource incompleteness, we show that existing lexica…