Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
arXiv:2206.04615
Abstract
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyond the Imitation Game benchmark (BIG-bench). BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions. Task topics are diverse, drawing problems from linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias, software development, and beyond. BIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models. We evaluate the behavior of OpenAI's GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters. In addition, a team of human expert raters performed all tasks in order to provide a strong baseline. Findings include: model performance and calibration both improve with scale, but are poor in absolute terms (and when compared with rater performance); performance is remarkably similar across model classes, though with benefits from sparsity; tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit "breakthrough" behavior at a critical scale often involve multiple steps or components, or brittle metrics; social bias typically increases with scale in settings with ambiguous context, but this can be improved with prompting.
27 pages, 17 figures + references and appendices, repo: https://github.com/google/BIG-bench
Cited by in corpus (52)
- ChatGPT: Jack of all trades, master of none
- Using cognitive psychology to understand GPT-3
- Unleashing the potential of prompt engineering for large language models
- Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
- 14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon
- What Large Language Models Know and What People Think They Know
- Factuality Challenges in the Era of Large Language Models
- Large language models surpass human experts in predicting neuroscience results
- Harms from Increasingly Agentic Algorithmic Systems
- OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
- Materials science in the era of large language models: a perspective
- GPT has become financially literate: Insights from financial literacy tests of GPT and a preliminary test of how people use it as a source of advice
- Large Language Models as Zero-Shot Human Models for Human-Robot Interaction
- State-of-the-art generalisation research in NLP: A taxonomy and review
- LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
- Hallucination Detection in Foundation Models for Decision-Making: A Flexible Definition and Review of the State of the Art
- CatAlyst: Domain-Extensible Intervention for Preventing Task Procrastination Using Large Generative Models
- Higher education assessment practice in the era of generative AI tools
- Exploring the Sensitivity of LLMs' Decision-Making Capabilities: Insights from Prompt Variation and Hyperparameters
- Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
- Towards The Ultimate Brain: Exploring Scientific Discovery with ChatGPT AI
- Introducing MBIB -- the first Media Bias Identification Benchmark Task and Dataset Collection
- Natural Language Processing with Commonsense Knowledge: A Survey
- KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts
- Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling
- Potential Benefits of Employing Large Language Models in Research in Moral Education and Development
- SmartLLMSentry: A Comprehensive LLM Based Smart Contract Vulnerability Detection Framework
- Quantum Many-Body Physics Calculations with Large Language Models
- GenQREnsemble: Zero-Shot LLM Ensemble Prompting for Generative Query Reformulation
- The Life Cycle of Knowledge in Big Language Models: A Survey
- On Sarcasm Detection with OpenAI GPT-based Models
- Transforming Agency. On the mode of existence of Large Language Models
- Large Language Models for Combinatorial Optimization: A Systematic Review
- GPT Struct Me: Probing GPT Models on Narrative Entity Extraction
- Complex QA and language models hybrid architectures, Survey
- Relational Programming with Foundation Models
- Piloting Copilot, Codex, and StarCoder2: Hot Temperature, Cold Prompts, or Black Magic?
- Structural Similarities Between Language Models and Neural Response Measurements
- SMILE: Evaluation and Domain Adaptation for Social Media Language Understanding
- Evaluating Language Model Agency through Negotiations
- The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
- LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
- Automated Theorem Provers Help Improve Large Language Model Reasoning
- Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case Study
- On the attribution of confidence to large language models
- Towards More Robust NLP System Evaluation: Handling Missing Scores in Benchmarks
- Equilibration of Coordinating Imitation and Best-Response Dynamics
- The Influence of Faulty Labels in Data Sets on Human Pose Estimation
- Do Large Language Models know who did what to whom?
- LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
- Searching Personal Collections
- Emergent inabilities? Inverse scaling over the course of pretraining