collaborators

7 papers

cs.CL2026

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala

Minuri Rajapakse, Ruvan Weerasinghe

The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly regarding script variation in digital…

cs.CL2026

Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources

Deshan Sumanathilaka, Sameera Perera, Sachithya Dharmasiri +6

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. T…

cs.CL2026

SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech

Johan Sofalas, Dilushri Pavithra, Nevidu Jayatilleke +1

Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative…

cs.CL2025

SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala

Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda +6

Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, of…

cs.CL2025

A Hybrid Architecture with Efficient Fine Tuning for Abstractive Patent Document Summarization

Nevidu Jayatilleke, Ruvan Weerasinghe

Automatic patent summarization approaches that help in the patent analysis and comprehension procedure are in high demand due to the colossal growth of innovations. The development…

cs.CL2025

Advancements in Natural Language Processing for Automatic Text Summarization

Nevidu Jayatilleke, Ruvan Weerasinghe, Nipuna Senanayake

The substantial growth of textual content in diverse domains and platforms has led to a considerable need for Automatic Text Summarization (ATS) techniques that aid in the process…