1 paper · 1 filter
Sergey Berezin, Reza Farahbakhsh, Noel Crespi
We introduce a novel class of adversarial attacks on toxicity detection models that exploit language models' failure to interpret spatially structured text in the form of ASCII art…