activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL202612 cited

Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive

Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe +3

Offensive speech detection is a key component of content moderation. However, what is offensive can be highly subjective. This paper investigates how machine and human moderators d…

cs.CL20251 cited

What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge

Arka Dutta, Sujan Dutta, Rijul Magu +3

Hallucinations pose a critical challenge to the real-world deployment of large language models (LLMs) in high-stakes domains. In this paper, we present a framework for stress testi…

cs.CL2024

Rater Cohesion and Quality from a Vicarious Perspective

Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta +5

Human feedback is essential for building human-centered AI systems across domains where disagreement is prevalent, such as AI safety, content moderation, or sentiment analysis. Man…

cs.CL2024

ARTICLE: Annotator Reliability Through In-Context Learning

Sujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya +3

Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsica…

cs.CL2024

Gender Representation and Bias in Indian Civil Service Mock Interviews

Somonnoy Banerjee, Sujan Dutta, Soumyajit Datta +1

This paper makes three key contributions. First, via a substantial corpus of 51,278 interview questions sourced from 888 YouTube videos of mock interviews of Indian civil service c…

cs.CL2024

Applying RLAIF for Code Generation with API-usage in Lightweight LLMs

Sujan Dutta, Sayantan Mahinder, Raviteja Anantha +1

Reinforcement Learning from AI Feedback (RLAIF) has demonstrated significant potential across various domains, including mitigating harm in LLM outputs, enhancing text summarizatio…