3 papers
cs.CL2025
Compositional Generalisation for Explainable Hate Speech Detection
Agostina Calabrese, Tom Sherborne, Björn Ross +1
Hate speech detection is key to online content moderation, but current models struggle to generalise beyond their training data. This has been linked to dataset biases and the use…
cs.CL2024
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster
Agostina Calabrese, Leonardo Neves, Neil Shah +4
Content moderators play a key role in keeping the conversation on social media healthy. While the high volume of content they need to judge represents a bottleneck to the moderatio…
cs.CL2022
Explainable Abuse Detection as Intent Classification and Slot Filling
Agostina Calabrese, Björn Ross, Mirella Lapata
To proactively offer social media users a safe online experience, there is a need for systems that can detect harmful posts and promptly alert platform moderators. In order to guar…