2 papers
cs.CL2024
Disentangling Hate Across Target Identities
Yiping Jin, Leo Wanner, Aneesh Moideen Koya
Hate speech (HS) classifiers do not perform equally well in detecting hateful expressions towards different target identities. They also demonstrate systematic biases in predicted…
cs.CL2024
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
Yiping Jin, Leo Wanner, Alexander Shvets
Online hate detection suffers from biases incurred in data sampling, annotation, and model pre-training. Therefore, measuring the averaged performance over all examples in held-out…