activity
20202022
most citedHate Speech Classifiers Learn Human-Like Social Stereotypes

4 citations · 7 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL20222 cited

Social-Group-Agnostic Word Embedding Debiasing via the Stereotype Content Model

Ali Omrani, Brendan Kennedy, Mohammad Atari +1

Existing word embedding debiasing methods require social-group-specific word pairs (e.g., "man"-"woman") for each social attribute (e.g., gender), which cannot be used to mitigate…

cs.CL20214 cited

Hate Speech Classifiers Learn Human-Like Social Stereotypes

Aida Mostafazadeh Davani, Mohammad Atari, Brendan Kennedy +1

Social stereotypes negatively impact individuals' judgements about different groups and may have a critical role in how people understand language directed toward minority social g…

cs.CL20211 cited

Improving Counterfactual Generation for Fair Hate Speech Detection

Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy +3

Bias mitigation approaches reduce models' dependence on sensitive features of data, such as social group tokens (SGTs), resulting in equal predictions across the sensitive features…

cs.CL2020

Fair Hate Speech Detection through Evaluation of Social Group Counterfactuals

Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy +3

Approaches for mitigating bias in supervised models are designed to reduce models' dependence on specific sensitive features of the input data, e.g., mentioned social groups. Howev…

cs.CL2020

On Transferability of Bias Mitigation Effects in Language Model Fine-Tuning

Xisen Jin, Francesco Barbieri, Brendan Kennedy +3

Fine-tuned language models have been shown to exhibit biases against protected groups in a host of modeling tasks such as text classification and coreference resolution. Previous w…