7 papers
Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala
Minuri Rajapakse, Ruvan Weerasinghe
The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly regarding script variation in digital…
Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources
Deshan Sumanathilaka, Sameera Perera, Sachithya Dharmasiri +6
The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. T…
SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech
Johan Sofalas, Dilushri Pavithra, Nevidu Jayatilleke +1
Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative…
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda +6
Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, of…
A Hybrid Architecture with Efficient Fine Tuning for Abstractive Patent Document Summarization
Nevidu Jayatilleke, Ruvan Weerasinghe
Automatic patent summarization approaches that help in the patent analysis and comprehension procedure are in high demand due to the colossal growth of innovations. The development…
Advancements in Natural Language Processing for Automatic Text Summarization
Nevidu Jayatilleke, Ruvan Weerasinghe, Nipuna Senanayake
The substantial growth of textual content in diverse domains and platforms has led to a considerable need for Automatic Text Summarization (ATS) techniques that aid in the process…