13 citations · 28 across the 21 of their papers we have counts for
16 papers · 1 filter
Dynamics of meaning: Towards the Evaluation of Diachronic Semantic Change in Sinhala
Nevidu Jayatilleke, Nisansa de Silva
Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and the limitations of static embeddin…
HelaBERT: Enhancing Sinhala Language Understanding with Dual Pooling Classification Head
Thisen Ekanayake, Nisansa de Silva
We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourced from MADLAD-400, CulturaX,…
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
Theekshana Samaradiwakara, Nisansa de Silva, George C. Lobb
Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural L…
SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis
Lakshani Galwatta, Nisansa de Silva, Sarangi Aththanayake +1
Sentiment analysis has been a primary domain under Natural Language Processing (NLP) from its inception as it plays a vital role in both real-world and research applications. In hi…
SiDiaC: Sinhala Diachronic Corpus
Nevidu Jayatilleke, Nisansa de Silva
SiDiaC, the first comprehensive Sinhala Diachronic Corpus, covers a historical span from the 5th to the 20th century CE. SiDiaC comprises 58k words across 46 literary works, annota…
SinLlama -- A Large Language Model for Sinhala
H. W. K. Aravinda, Rashad Sirajudeen, Samith Karunathilake +3
Low-resource languages such as Sinhala are often overlooked by open-source Large Language Models (LLMs). In this research, we extend an existing multilingual LLM (Llama-3-8B) to be…