activity
20202025
most citedData Augmentation and Terminology Integration for Domain-Specific Sinhala-English-Tamil Statistical Machine Translation

13 citations · 20 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2025

Improving the quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics

Aloka Fernando, Nisansa de Silva, Menan Velyuthan +2

Parallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from web-mined corpora. Ranking sentence pairs using similarity scores on sentence embeddings der…

cs.CL2025★ 1 cited

Linguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource Languages

Aloka Fernando, Surangika Ranathunga

Multilingual Pre-trained Language models (multiPLMs), trained on the Masked Language Modelling (MLM) objective are commonly being used for cross-lingual tasks such as bitext mining…

cs.CL2024

Shoulders of Giants: A Look at the Degree and Utility of Openness in NLP Research

Surangika Ranathunga, Nisansa de Silva, Dilith Jayakody +1

We analysed a sample of NLP research papers archived in ACL Anthology as an attempt to quantify the degree of openness and the benefit of such an open culture in the NLP community.…

cs.CL2024

Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora

Surangika Ranathunga, Nisansa de Silva, Menan Velayuthan +2

We conducted a detailed analysis on the quality of web-mined corpora for two low-resource languages (making three language pairs, English-Sinhala, English-Tamil and Sinhala-Tamil).…

cs.CL2022★ 6 cited

Data Augmentation to Address Out-of-Vocabulary Problem in Low-Resource Sinhala-English Neural Machine Translation

Aloka Fernando, Surangika Ranathunga

Out-of-Vocabulary (OOV) is a problem for Neural Machine Translation (NMT). OOV refers to words with a low occurrence in the training data, or to those that are absent from the trai…

cs.CL2020★ 13 cited

Data Augmentation and Terminology Integration for Domain-Specific Sinhala-English-Tamil Statistical Machine Translation

Aloka Fernando, Surangika Ranathunga, Gihan Dias

Out of vocabulary (OOV) is a problem in the context of Machine Translation (MT) in low-resourced languages. When source and/or target languages are morphologically rich, it becomes…