20 citations · 25 across the 3 of their papers we have counts for
6 papers
Hate Speech Classifiers Learn Human-Like Social Stereotypes
Aida Mostafazadeh Davani, Mohammad Atari, Brendan Kennedy +1
Social stereotypes negatively impact individuals' judgements about different groups and may have a critical role in how people understand language directed toward minority social g…
Improving Counterfactual Generation for Fair Hate Speech Detection
Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy +3
Bias mitigation approaches reduce models' dependence on sensitive features of data, such as social group tokens (SGTs), resulting in equal predictions across the sensitive features…
Fair Hate Speech Detection through Evaluation of Social Group Counterfactuals
Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy +3
Approaches for mitigating bias in supervised models are designed to reduce models' dependence on specific sensitive features of the input data, e.g., mentioned social groups. Howev…
On Transferability of Bias Mitigation Effects in Language Model Fine-Tuning
Xisen Jin, Francesco Barbieri, Brendan Kennedy +3
Fine-tuned language models have been shown to exhibit biases against protected groups in a host of modeling tasks such as text classification and coreference resolution. Previous w…
Contextualizing Hate Speech Classifiers with Post-hoc Explanation
Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani +2
Hate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways. Such biases manife…
Reporting the Unreported: Event Extraction for Analyzing the Local Representation of Hate Crimes
Aida Mostafazadeh Davani, Leigh Yeh, Mohammad Atari +8
Official reports of hate crimes in the US are under-reported relative to the actual number of such incidents. Further, despite statistical approximations, there are no official rep…