3 papers
cs.CL2025
Characterizing Selective Refusal Bias in Large Language Models
Adel Khorramrouz, Sharon Levy
Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently…
cs.CL2023
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
Arka Dutta, Adel Khorramrouz, Sujan Dutta +1
This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide su…
cs.CY2023
For Women, Life, Freedom: A Participatory AI-Based Social Web Analysis of a Watershed Moment in Iran's Gender Struggles
Adel Khorramrouz, Sujan Dutta, Ashiqur R. KhudaBukhsh
In this paper, we present a computational analysis of the Persian language Twitter discourse with the aim to estimate the shift in stance toward gender equality following the death…