Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Characterizing Selective Refusal Bias in Large Language Models
Adel Khorramrouz, Sharon Levy
Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently…
cs.CL2024
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
Arka Dutta, Adel Khorramrouz, Sujan Dutta +1
This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide su…