2 papers
cs.CL2025
Compositional Generalisation for Explainable Hate Speech Detection
Agostina Calabrese, Tom Sherborne, Björn Ross +1
Hate speech detection is key to online content moderation, but current models struggle to generalise beyond their training data. This has been linked to dataset biases and the use…
cs.CL2024
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster
Agostina Calabrese, Leonardo Neves, Neil Shah +4
Content moderators play a key role in keeping the conversation on social media healthy. While the high volume of content they need to judge represents a bottleneck to the moderatio…