Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
arXiv:2401.12295 · doi:10.1177/00491241251340608
Abstract
The field of machine learning has recently made significant progress in reducing the requirements for labelled training data when building new models. These `cheaper' learning techniques hold significant potential for the social sciences, where development of large labelled training datasets is often a significant practical impediment to the use of machine learning for analytical tasks. In this article we review three `cheap' techniques that have developed in recent years: weak supervision, transfer learning and prompt engineering. For the latter, we also review the particular case of zero-shot prompting of large language models. For each technique we provide a guide of how it works and demonstrate its application across six different realistic social science applications (two different tasks paired with three different dataset makeups). We show good performance for all techniques, and in particular we demonstrate how prompting of large language models can achieve high accuracy at very low cost. Our results are accompanied by a code repository to make it easy for others to duplicate our work and use it in their own research. Overall, our article is intended to stimulate further uptake of these techniques in the social sciences.
46 pages, 17 figures, 6 tables
References in corpus (25)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Cross-lingual Language Model Pretraining
- Large Language Models are Zero-Shot Reasoners
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- Snorkel: Rapid Training Data Creation with Weak Supervision
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Data Programming: Creating Large Training Sets, Quickly
- What is being transferred in transfer learning?
- Parameter-Efficient Transfer Learning for NLP
- Ontology-driven weak supervision for clinical entity classification in electronic health records
- Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
- Finetuned Language Models Are Zero-Shot Learners
- A Mathematical Exploration of Why Language Models Help Solve Downstream Tasks
- Towards Efficient Post-training Quantization of Pre-trained Language Models
- Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
- Open, Closed, or Small Language Models for Text Classification?
- Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT
- Semi-supervised Classification for Natural Language Processing
- High-dimensional Imputation for the Social Sciences: a Comparison of State-of-the-art Methods
- Customer Sentiment Analysis using Weak Supervision for Customer-Agent Chat
- Drawing Causal Inferences About Performance Effects in NLP