13 citations · 27 across the 20 of their papers we have counts for
18 papers · 1 filter
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates
Himath Dhanapala, Haren Daishika, Himandhi Kuruppu +6
Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to stan…
LKValues: Aligning Large Language Models with Sri Lankan Societal Values
Nethmi Muthugala, Supryadi, Surangika Ranathunga +7
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societi…
Fault of Our Stars: Behavioral Drivers of Rating-Sentiment Incongruence
Ramanaish Abaiyan, Ruththiragayan Sutharsan, Kusal Amantha +7
When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely used as convenient weak…
Sinhala Physical Common Sense Reasoning Dataset for Global PIQA
Nisansa de Silva, Surangika Ranathunga
This paper presents the first-ever Sinhala physical common sense reasoning dataset created as part of Global PIQA. It contains 110 human-created and verified data samples, where ea…
LMSpell: Neural Spell Checking for Low-Resource Languages
Akesh Gunathilake, Nadil Karunarathna, Tharusha Bandaranayake +2
Spell correction is still a challenging problem for low-resource languages (LRLs). While pretrained language models (PLMs) have been employed for spell correction, their use is sti…
GeeSanBhava: Sentiment Tagged Sinhala Music Video Comment Data Set
Yomal De Mel, Nisansa de Silva
This study introduce GeeSanBhava, a high-quality data set of Sinhala song comments extracted from YouTube manually tagged using Russells Valence-Arousal model by three independent…