13 citations · 14 across the 3 of their papers we have counts for
8 papers
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia
Alham Fikri Aji, Genta Indra Winata, Fajri Koto +9
NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia,…
Unsupervised Cross-Lingual Transfer of Structured Predictors without Source Data
Kemal Kurniawan, Lea Frermann, Philip Schulz +1
Providing technologies to communities or domains where training data is scarce or protected e.g., for privacy reasons, is becoming increasingly important. To that end, we generalis…
PPT: Parsimonious Parser Transfer for Unsupervised Cross-Lingual Adaptation
Kemal Kurniawan, Lea Frermann, Philip Schulz +1
Cross-lingual transfer is a leading technique for parsing low-resource languages in the absence of explicit supervision. Simple `direct transfer' of a learned model based on a mult…
KaWAT: A Word Analogy Task Dataset for Indonesian
Kemal Kurniawan
We introduced KaWAT (Kata Word Analogy Task), a new word analogy task dataset for Indonesian. We evaluated on it several existing pretrained Indonesian word embeddings and embeddin…
IndoSum: A New Benchmark Dataset for Indonesian Text Summarization
Kemal Kurniawan, Samuel Louvan
Automatic text summarization is generally considered as a challenging task in the NLP community. One of the challenges is the publicly available and large dataset that is relativel…
Toward a Standardized and More Accurate Indonesian Part-of-Speech Tagging
Kemal Kurniawan, Alham Fikri Aji
Previous work in Indonesian part-of-speech (POS) tagging are hard to compare as they are not evaluated on a common dataset. Furthermore, in spite of the success of neural network m…