activity
20202022
most citedContextualizing Hate Speech Classifiers with Post-hoc Explanation

20 citations · 27 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL20222 cited

Social-Group-Agnostic Word Embedding Debiasing via the Stereotype Content Model

Ali Omrani, Brendan Kennedy, Mohammad Atari +1

Existing word embedding debiasing methods require social-group-specific word pairs (e.g., "man"-"woman") for each social attribute (e.g., gender), which cannot be used to mitigate…

cs.CL20214 cited

Hate Speech Classifiers Learn Human-Like Social Stereotypes

Aida Mostafazadeh Davani, Mohammad Atari, Brendan Kennedy +1

Social stereotypes negatively impact individuals' judgements about different groups and may have a critical role in how people understand language directed toward minority social g…

cs.CL20211 cited

Improving Counterfactual Generation for Fair Hate Speech Detection

Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy +3

Bias mitigation approaches reduce models' dependence on sensitive features of data, such as social group tokens (SGTs), resulting in equal predictions across the sensitive features…

cs.CL2020

Fair Hate Speech Detection through Evaluation of Social Group Counterfactuals

Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy +3

Approaches for mitigating bias in supervised models are designed to reduce models' dependence on specific sensitive features of the input data, e.g., mentioned social groups. Howev…

cs.CL202020 cited

Contextualizing Hate Speech Classifiers with Post-hoc Explanation

Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani +2

Hate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways. Such biases manife…