1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2023
No offence, Bert -- I insult only humans! Multiple addressees sentence-level attack on toxicity detection neural network
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
We introduce a simple yet efficient sentence-level attack on black-box toxicity detector models. By adding several positive words or sentences to the end of a hateful message, we a…
cs.CL2023★ 1 cited
On the definition of toxicity in NLP
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
The fundamental problem in toxicity detection task lies in the fact that the toxicity is ill-defined. This causes us to rely on subjective and vague data in models' training, which…