A Large Self-Annotated Corpus for Sarcasm
arXiv:1704.05579
Abstract
We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous dataset -- and many times more instances of non-sarcastic statements, allowing for learning in both balanced and unbalanced label regimes. Each statement is furthermore self-annotated -- sarcasm is labeled by the author, not an independent annotator -- and provided with user, topic, and conversation context. We evaluate the corpus for accuracy, construct benchmarks for sarcasm detection, and evaluate baseline methods.
6 pages, 4 Figures. To Appear in LREC 2018
References in corpus (2)
Cited by in corpus (11)
- A Transformer-based approach to Irony and Sarcasm detection
- CASCADE: Contextual Sarcasm Detection in Online Discussion Forums
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
- "President Vows to Cut <Taxes> Hair": Dataset and Analysis of Creative Text Editing for Humorous Headlines
- Reasoning with Sarcasm by Reading In-between
- Reddit Entity Linking Dataset
- A Survey of Multimodal Sarcasm Detection
- Do Humans Trust Advice More if it Comes from AI? An Analysis of Human-AI Interactions
- On Sarcasm Detection with OpenAI GPT-based Models
- Parallel Deep Learning-Driven Sarcasm Detection from Pop Culture Text and English Humor Literature
- Representing Social Media Users for Sarcasm Detection