1 citations · 1 across the 5 of their papers we have counts for
3 papers · 1 filter
Evading Toxicity Detection with ASCII-art: A Benchmark of Spatial Attacks on Moderation Systems
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
We introduce a novel class of adversarial attacks on toxicity detection models that exploit language models' failure to interpret spatially structured text in the form of ASCII art…
No offence, Bert -- I insult only humans! Multiple addressees sentence-level attack on toxicity detection neural network
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
We introduce a simple yet efficient sentence-level attack on black-box toxicity detector models. By adding several positive words or sentences to the end of a hateful message, we a…
On the definition of toxicity in NLP
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
The fundamental problem in toxicity detection task lies in the fact that the toxicity is ill-defined. This causes us to rely on subjective and vague data in models' training, which…