37 citations · 57 across the 9 of their papers we have counts for
16 papers
MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models
David Setiawan, Temuulen Khishigsuren, Milind Agarwal +3
Multilingual dictionaries are among the most valuable documentary resources for low-resource and endangered languages, yet many remain available only as scans. For many decades, th…
CommonMorph: Participatory Morphological Documentation Platform
Aso Mahmudi, Sina Ahmadi, Kemal Kurniawan +3
Collecting and annotating morphological data present significant challenges, requiring linguistic expertise, methodological rigour, and substantial resources. These barriers are pa…
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers +4
The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance…
UniMorph 4.0: Universal Morphology
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa +93
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world…
The SIGMORPHON 2022 Shared Task on Morpheme Segmentation
Khuyagbaatar Batsuren, Gábor Bella, Aryaman Arora +10
The SIGMORPHON 2022 shared task on morpheme segmentation challenged systems to decompose a word into a sequence of morphemes and covered most types of morphology: compounds, deriva…
SIGTYP 2021 Shared Task: Robust Spoken Language Identification
Elizabeth Salesky, Badr M. Abdullah, Sabrina J. Mielke +6
While language identification is a fundamental speech and language processing task, for many languages and language families it remains a challenging task. For many low-resource an…