ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
arXiv:2303.15056 · doi:10.1073/pnas.2305016120
Abstract
Many NLP applications require manual data annotations for a variety of tasks, notably to train classifiers or evaluate the performance of unsupervised models. Depending on the size and degree of complexity, the tasks may be conducted by crowd-workers on platforms such as MTurk as well as trained annotators, such as research assistants. Using a sample of 2,382 tweets, we demonstrate that ChatGPT outperforms crowd-workers for several annotation tasks, including relevance, stance, topics, and frames detection. Specifically, the zero-shot accuracy of ChatGPT exceeds that of crowd-workers for four out of five tasks, while ChatGPT's intercoder agreement exceeds that of both crowd-workers and trained annotators for all tasks. Moreover, the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk. These results show the potential of large language models to drastically increase the efficiency of text classification.
Gilardi, Fabrizio, Meysam Alizadeh, and Maël Kubli. 2023. "ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks". Proceedings of the National Academy of Sciences 120(30): e2305016120
References in corpus (4)
Cited by in corpus (62)
- Generative AI
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
- A Scoping Review of ChatGPT Research in Accounting and Finance
- A Large Language Model Approach to Educational Survey Feedback Analysis
- The illusion of artificial inclusion
- Quilt-1M: One Million Image-Text Pairs for Histopathology
- Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
- The simulation of judgment in LLMs
- Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation
- Automated stance detection in complex topics and small languages: the challenging case of immigration in polarizing news media
- Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges
- Automated Review Generation Method Based on Large Language Models
- Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
- Stance Detection: A Practical Guide to Classifying Political Beliefs in Text
- AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
- Doing Personal LAPS: LLM-Augmented Dialogue Construction for Personalized Multi-Session Conversational Search
- Automated Assessment of Encouragement and Warmth in Classrooms Leveraging Multimodal Emotional Features and ChatGPT
- Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts
- Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
- Can GPT-4 learn to analyse moves in research article abstracts?
- Query Performance Prediction using Relevance Judgments Generated by Large Language Models
- An Empirical Study of Challenges in Machine Learning Asset Management
- Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
- Laboratory-Scale AI: Open-Weight Models are Competitive with ChatGPT Even in Low-Resource Settings
- Framing Social Movements on Social Media: Unpacking Diagnostic, Prognostic, and Motivational Strategies
- Health-promoting Potential of Parks in 35 Cities Worldwide
- Human-interpretable clustering of short-text using large language models
- Annotation Guidelines-Based Knowledge Augmentation: Towards Enhancing Large Language Models for Educational Text Classification
- Discrete Prompt Compression with Reinforcement Learning
- Artificial Intelligence Can Emulate Human Normative Judgments on Emotional Visual Scenes
- APT-Pipe: A Prompt-Tuning Tool for Social Data Annotation using ChatGPT
- Leveraging Prompt-Based Large Language Models: Predicting Pandemic Health Decisions and Outcomes Through Social Media Language
- Synthetically generated text for supervised text analysis
- From Voices to Validity: Leveraging Large Language Models (LLMs) for Textual Analysis of Policy Stakeholder Interviews
- Judgment of Learning: A Human Ability Beyond Generative Artificial Intelligence
- GPT Assisted Annotation of Rhetorical and Linguistic Features for Interpretable Propaganda Technique Detection in News Text
- Text-to-SQL Domain Adaptation via Human-LLM Collaborative Data Annotation
- Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models
- FAIL: Analyzing Software Failures from the News Using LLMs
- Using language models to label clusters of scientific documents
- SpreadLine: Visualizing Egocentric Dynamic Influence
- Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text
- Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
- Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
- Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins
- Assessing Good, Bad and Ugly Arguments Generated by ChatGPT: a New Dataset, its Methodology and Associated Tasks
- VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data
- Infrastructure Ombudsman: Mining Future Failure Concerns from Structural Disaster Response
- Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
- Human Interest or Conflict? Leveraging LLMs for Automated Framing Analysis in TV Shows
- Exploring news intent and its application: A theory-driven approach
- Large Language Models Are Democracy Coders with Attitudes
- Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
- Semantically Orthogonal Framework for Citation Classification: Disentangling Intent and Content
- Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
- BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts
- Embracing Dialectic Intersubjectivity: Coordination of Different Perspectives in Content Analysis with LLM Persona Simulation
- Towards a Psychology of Machines: Large Language Models Predict Human Memory
- Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications