9 papers
From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition
Lukmal Ilyas, Nevidu Jayatilleke
Dhivehi, the national language of the Maldives, is currently under-resourced for automatic speech recognition (ASR) and other NLP tasks. This study investigates whether cross-lingu…
Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis
Avisha Dilhara, Nevidu Jayatilleke
Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala…
SiPaKosa: A Comprehensive Corpus of Canonical and Classical Buddhist Texts in Sinhala and Pali
Ranidu Gurusinghe, Nevidu Jayatilleke
SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhi…
SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi +3
SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication dates, and a historical span from the…
SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech
Johan Sofalas, Dilushri Pavithra, Nevidu Jayatilleke +1
Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative…
SiDiaC: Sinhala Diachronic Corpus
Nevidu Jayatilleke, Nisansa de Silva
SiDiaC, the first comprehensive Sinhala Diachronic Corpus, covers a historical span from the 5th to the 20th century CE. SiDiaC comprises 58k words across 46 literary works, annota…