Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities
arXiv:2406.07353 · doi:10.1016/j.osnem.2025.100317
Abstract
Internet memes, channels for humor, social commentary, and cultural expression, are increasingly used to spread toxic messages. Studies on the computational analyses of toxic memes have significantly grown over the past five years, and the only three surveys on computational toxic meme analysis cover only work published until 2022, leading to inconsistent terminology and unexplored trends. Our work fills this gap by surveying content-based computational perspectives on toxic memes, and reviewing key developments until early 2024. Employing the PRISMA methodology, we systematically extend the previously considered papers, achieving a threefold result. First, we survey 119 new papers, analyzing 158 computational works focused on content-based toxic meme analysis. We identify over 30 datasets used in toxic meme analysis and examine their labeling systems. Second, after observing the existence of unclear definitions of meme toxicity in computational works, we introduce a new taxonomy for categorizing meme toxicity types. We also note an expansion in computational tasks beyond the simple binary classification of memes as toxic or non-toxic, indicating a shift towards achieving a nuanced comprehension of toxicity. Third, we identify three content-based dimensions of meme toxicity under automatic study: target, intent, and conveyance tactics. We develop a framework illustrating the relationships between these dimensions and meme toxicities. The survey analyzes key challenges and recent trends, such as enhanced cross-modal reasoning, integrating expert and cultural knowledge, the demand for automatic toxicity explanations, and handling meme toxicity in low-resource languages. Also, it notes the rising use of Large Language Models (LLMs) and generative AI for detecting and generating toxic memes. Finally, it proposes pathways for advancing toxic meme detection and interpretation.
39 pages, 12 figures, 9 tables
References in corpus (34)
- Exploring Lightweight Interventions at Posting Time to Reduce the Sharing of Misinformation on Social Media
- A Framework of Severity for Harmful Content Online
- Defining and Detecting Toxicity on Social Media: Context and Knowledge are Key
- Disentangling Hate in Online Memes
- MIntRec: A New Dataset for Multimodal Intent Recognition
- Hate Speech in Pixels: Detection of Offensive Memes towards Automatic Moderation
- Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge
- A Multimodal Framework for the Detection of Hateful Memes
- Enhance Multimodal Transformer With External Label And In-Domain Pretrain: Hateful Meme Challenge Winning Solution
- Detecting Hate Speech in Multi-modal Memes
- Vilio: State-of-the-art Visio-Linguistic Models applied to Hateful Memes
- Benchmark dataset of memes with text transcriptions for automatic detection of multi-modal misogynistic content
- Detecting Hateful Memes Using a Multimodal Deep Ensemble
- MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
- Hateful Memes Detection via Complementary Visual and Linguistic Networks
- Classification of Multimodal Hate Speech -- The Winning Solution of Hateful Memes Challenge
- Do Images really do the Talking? Analysing the significance of Images in Tamil Troll meme classification
- DisinfoMeme: A Multimodal Dataset for Detecting Meme Intentionally Spreading Out Disinformation
- BLUE at Memotion 2.0 2022: You have my Image, my Text and my Transformer
- Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
- Feels Bad Man: Dissecting Automated Hateful Meme Detection Through the Lens of Facebook's Challenge
- GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse
- On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
- Hate Me Not: Detecting Hate Inducing Memes in Code Switched Languages
- Detection of Propaganda Techniques in Visuo-Lingual Metaphor in Memes
- Detecting and Correcting Hate Speech in Multimodal Memes with Large Visual Language Model
- Text or Image? What is More Important in Cross-Domain Generalization Capabilities of Hate Meme Detection Models?
- Hateful Memes Challenge: An Enhanced Multimodal Framework
- Enhance Multimodal Model Performance with Data Augmentation: Facebook Hateful Meme Challenge Solution
- A Review of Vision-Language Models and their Performance on the Hateful Memes Challenge
- Causal Intersectionality and Dual Form of Gradient Descent for Multimodal Analysis: a Case Study on Hateful Memes
- Meme-ingful Analysis: Enhanced Understanding of Cyberbullying in Memes Through Multimodal Explanations
- The Hateful Memes Challenge Next Move
- A Template Is All You Meme