22 papers · 1 filter
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates
Himath Dhanapala, Haren Daishika, Himandhi Kuruppu +6
Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to stan…
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
Theekshana Samaradiwakara, Nisansa de Silva, George C. Lobb
Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural L…
LKValues: Aligning Large Language Models with Sri Lankan Societal Values
Nethmi Muthugala, Supryadi, Surangika Ranathunga +7
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societi…
SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis
Lakshani Galwatta, Nisansa de Silva, Sarangi Aththanayake +1
Sentiment analysis has been a primary domain under Natural Language Processing (NLP) from its inception as it plays a vital role in both real-world and research applications. In hi…
Fault of Our Stars: Behavioral Drivers of Rating-Sentiment Incongruence
Ramanaish Abaiyan, Ruththiragayan Sutharsan, Kusal Amantha +7
When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely used as convenient weak…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…