HateProof: Are Hateful Meme Detection Systems really Robust?
arXiv:2302.05703 · doi:10.1145/3543507.3583356
Abstract
Exploiting social media to spread hate has tremendously increased over the years. Lately, multi-modal hateful content such as memes has drawn relatively more traction than uni-modal content. Moreover, the availability of implicit content payloads makes them fairly challenging to be detected by existing hateful meme detection systems. In this paper, we present a use case study to analyze such systems' vulnerabilities against external adversarial attacks. We find that even very simple perturbations in uni-modal and multi-modal settings performed by humans with little knowledge about the model can make the existing detection models highly vulnerable. Empirically, we find a noticeable performance drop of as high as 10% in the macro-F1 score for certain attacks. As a remedy, we attempt to boost the model's robustness using contrastive learning as well as an adversarial training-based method - VILLA. Using an ensemble of the above two approaches, in two of our high resolution datasets, we are able to (re)gain back the performance to a large extent for certain attacks. We believe that ours is a first step toward addressing this crucial problem in an adversarial setting and would inspire more such investigations in the future.
Accepted at TheWebConf'2023 (WWW'2023)
References in corpus (10)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Theoretically Principled Trade-off between Robustness and Accuracy
- Hate Speech in Pixels: Detection of Offensive Memes towards Automatic Moderation
- Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge
- Vilio: State-of-the-art Visio-Linguistic Models applied to Hateful Memes
- A Closer Look at the Robustness of Vision-and-Language Pre-trained Models
- AugLy: Data Augmentations for Robustness
- Robustness Disparities in Commercial Face Detection
- Understanding and Measuring Robustness of Multimodal Learning
- On Explaining Multimodal Hateful Meme Detection Models