3 citations · 4 across the 3 of their papers we have counts for
3 papers
No offence, Bert -- I insult only humans! Multiple addressees sentence-level attack on toxicity detection neural network
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
We introduce a simple yet efficient sentence-level attack on black-box toxicity detector models. By adding several positive words or sentences to the end of a hateful message, we a…
On the definition of toxicity in NLP
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
The fundamental problem in toxicity detection task lies in the fact that the toxicity is ill-defined. This causes us to rely on subjective and vague data in models' training, which…
Hate Speech and Offensive Language Detection using an Emotion-aware Shared Encoder
Khouloud Mnassri, Praboda Rajapaksha, Reza Farahbakhsh +1
The rise of emergence of social media platforms has fundamentally altered how people communicate, and among the results of these developments is an increase in online use of abusiv…