3 papers
cs.CL2026
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe +3
Offensive speech detection is a key component of content moderation. However, what is offensive can be highly subjective. This paper investigates how machine and human moderators d…
cs.CL2025
What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge
Arka Dutta, Sujan Dutta, Rijul Magu +3
Hallucinations pose a critical challenge to the real-world deployment of large language models (LLMs) in high-stakes domains. In this paper, we present a framework for stress testi…
cs.CL2024
Rater Cohesion and Quality from a Vicarious Perspective
Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta +5
Human feedback is essential for building human-centered AI systems across domains where disagreement is prevalent, such as AI safety, content moderation, or sentiment analysis. Man…