2 papers
cs.LG2024
A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
Ivo Verhoeven, Pushkar Mishra, Rahel Beloch +2
Community models for malicious content detection, which take into account the context from a social graph alongside the content itself, have shown remarkable performance on benchma…
cs.CL2023
Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model
Abhijith Chintam, Rahel Beloch, Willem Zuidema +2
Language models (LMs) exhibit and amplify many types of undesirable biases learned from the training data, including gender bias. However, we lack tools for effectively and efficie…