PaLM: Scaling Language Modeling with Pathways
arXiv:2204.02311
Abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Cited by in corpus (158)
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- DINOv2: Learning Robust Visual Features without Supervision
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Recommender Systems in the Era of Large Language Models (LLMs)
- Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming
- Unleashing the potential of prompt engineering for large language models
- Compute Trends Across Three Eras of Machine Learning
- MetaFormer Baselines for Vision
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Supporting Qualitative Analysis with Large Language Models: Combining Codebook with GPT-3 for Deductive Coding
- Auditing large language models: a three-layered approach
- 14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon
- Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- Uncovering ChatGPT's Capabilities in Recommender Systems
- Tiny Machine Learning: Progress and Futures
- Large Language Models and the Reverse Turing Test
- AI-driven multi-omics integration for multi-scale predictive modeling of causal genotype-environment-phenotype relationships
- Evaluating Large Language Models on a Highly-specialized Topic, Radiation Oncology Physics
- ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
- Large Language Models for Code: Security Hardening and Adversarial Testing
- Large Language Models as Zero-Shot Conversational Recommenders
- Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers
- Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education
- Platform-Independent and Curriculum-Oriented Intelligent Assistant for Higher Education
- Grammatical Error Correction: A Survey of the State of the Art
- LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation
- OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- FAIR for AI: An interdisciplinary and international community building perspective
- LLM for SoC Security: A Paradigm Shift
- Mathemyths: Leveraging Large Language Models to Teach Mathematical Language through Child-AI Co-Creative Storytelling
- "HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
- Large Language Models as Zero-Shot Human Models for Human-Robot Interaction
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- LLM-based policy generation for intent-based management of applications
- Addressing Bias in Generative AI: Challenges and Research Opportunities in Information Management
- Theory of Mind for Multi-Agent Collaboration via Large Language Models
- Leveraging Large Language Models to Power Chatbots for Collecting User Self-Reported Data
- Understanding the Impact of Long-Term Memory on Self-Disclosure with Large Language Model-Driven Chatbots for Public Health Intervention
- Real-World Robot Applications of Foundation Models: A Review
- "We Need Structured Output": Towards User-centered Constraints on Large Language Model Output
- Refactoring Programs Using Large Language Models with Few-Shot Examples
- ChaCha: Leveraging Large Language Models to Prompt Children to Share Their Emotions about Personal Events
- Empowering Molecule Discovery for Molecule-Caption Translation with Large Language Models: A ChatGPT Perspective
- A Practical Survey on Zero-shot Prompt Design for In-context Learning
- Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
- AppPoet: Large Language Model based Android malware detection via multi-view prompt engineering
- HPC-GPT: Integrating Large Language Model for High-Performance Computing
- LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
- How Good is ChatGPT at Face Biometrics? A First Look into Recognition, Soft Biometrics, and Explainability
- MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
- Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model
- Embedding Large Language Models into Extended Reality: Opportunities and Challenges for Inclusion, Engagement, and Privacy
- Generative Relevance Feedback with Large Language Models
- Fine-tuning and Utilization Methods of Domain-specific LLMs
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- Exploring the Roles of Large Language Models in Reshaping Transportation Systems: A Survey, Framework, and Roadmap
- Enhanced Automated Code Vulnerability Repair using Large Language Models
- Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
- Leveraging Large Language Models for Patient Engagement: The Power of Conversational AI in Digital Health
- Let Me Do It For You: Towards LLM Empowered Recommendation via Tool Learning
- GalleryGPT: Analyzing Paintings with Large Multimodal Models
- Several categories of Large Language Models (LLMs): A Short Survey
- An Evaluation on Large Language Model Outputs: Discourse and Memorization
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- The Future of Intelligent Healthcare: A Systematic Analysis and Discussion on the Integration and Impact of Robots Using Large Language Models for Healthcare
- QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
- Exploiting Simulated User Feedback for Conversational Search: Ranking, Rewriting, and Beyond
- A dataset and benchmark for hospital course summarization with adapted large language models
- IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
- Prompted LLMs as Chatbot Modules for Long Open-domain Conversation
- Higher education assessment practice in the era of generative AI tools
- Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
- Revisiting the Plastic Surgery Hypothesis via Large Language Models
- A Framework for Neurosymbolic Robot Action Planning using Large Language Models
- Domain Terminology Integration into Machine Translation: Leveraging Large Language Models
- DR.BENCH: Diagnostic Reasoning Benchmark for Clinical Natural Language Processing
- Feedback-Driven Automated Whole Bug Report Reproduction for Android Apps
- VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
- A Study of Vulnerability Repair in JavaScript Programs with Large Language Models
- Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
- Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling
- Solving Math Word Problems via Cooperative Reasoning induced Language Models
- Seven Useful Questions in Density Functional Theory
- Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
- Exploring the Landscape of Ubiquitous In-home Health Monitoring: A Comprehensive Survey
- Supersonic: Learning to Generate Source Code Optimizations in C/C++
- Deep learning for nano-photonic materials -- The solution to everything!?
- Guideline Learning for In-context Information Extraction
- A Status Quo Investigation of Large Language Models towards Cost-Effective CFD Automation with OpenFOAMGPT: ChatGPT vs. Qwen vs. Deepseek
- Negative Human Rights as a Basis for Long-term AI Safety and Regulation
- Tucano: Advancing Neural Text Generation for Portuguese
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
- XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages
- Entity Recognition from Colloquial Text
- SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
- Towards Foundation Models for Materials Science: The Open MatSci ML Toolkit
- LAMBDA: A Large Model Based Data Agent
- CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM
- Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
- Towards Pareto Optimal Throughput in Small Language Model Serving
- Clinical information extraction for Low-resource languages with Few-shot learning using Pre-trained language models and Prompting
- Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models
- Measuring the Impact of Explanation Bias: A Study of Natural Language Justifications for Recommender Systems
- Symbolic Equation Solving via Reinforcement Learning
- Modularity in Deep Learning: A Survey
- A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
- GPT Struct Me: Probing GPT Models on Narrative Entity Extraction
- Towards Incremental Learning in Large Language Models: A Critical Review
- Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond
- Language Models for German Text Simplification: Overcoming Parallel Data Scarcity through Style-specific Pre-training
- Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data
- Differentiable Retrieval Augmentation via Generative Language Modeling for E-commerce Query Intent Classification
- Complex QA and language models hybrid architectures, Survey
- Piloting Copilot, Codex, and StarCoder2: Hot Temperature, Cold Prompts, or Black Magic?
- Developmental Scaffolding with Large Language Models
- Efficient Tuning of Large Language Models for Knowledge-Grounded Dialogue Generation
- CoqPyt: Proof Navigation in Python in the Era of LLMs
- ExClaim: Explainable Neural Claim Verification Using Rationalization
- Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
- MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
- Uncertainty Quantification and Decomposition for LLM-based Recommendation
- OPT-R: Exploring the Role of Explanations in Finetuning and Prompting for Reasoning Skills of Large Language Models
- Blacks is to Anger as Whites is to Joy? Understanding Latent Affective Bias in Large Pre-trained Neural Language Models
- SMILE: Evaluation and Domain Adaptation for Social Media Language Understanding
- Exploring the Impact of Model Scaling on Parameter-Efficient Tuning
- MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language Models
- ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
- EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
- PartIR: Composing SPMD Partitioning Strategies for Machine Learning
- A Mathematical Framework for Learning Probability Distributions
- Predictive Authoring for Brazilian Portuguese Augmentative and Alternative Communication
- AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database
- An LLM-based Agentic Framework for Accessible Network Control
- RAMP: Retrieval and Attribute-Marking Enhanced Prompting for Attribute-Controlled Translation
- Long-Range Transformer Architectures for Document Understanding
- On minimal variations for unsupervised representation learning
- LARR: Large Language Model Aided Real-time Scene Recommendation with Semantic Understanding
- Comprehensive Deadlock Prevention for GPU Collective Communication
- What is Wrong with Language Models that Can Not Tell a Story?
- MAX: Masked Autoencoder for X-ray Fluorescence in Geological Investigation
- Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization
- Model Fusion via Neuron Transplantation
- An Overview on Generative AI at Scale with Edge-Cloud Computing
- QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
- Linguacodus: A Synergistic Framework for Transformative Code Generation in Machine Learning Pipelines
- Tensorization of neural networks for improved privacy and interpretability
- CFL: Causally Fair Language Models Through Token-level Attribute Controlled Generation
- Leveraging Cross-Utterance Context For ASR Decoding
- It Ain't That Bad: Understanding the Mysterious Performance Drop in OOD Generalization for Generative Transformer Models
- Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
- Remember what you did so you know what to do next
- Program Skeletons for Automated Program Translation
- Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
- Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
- Simple and Effective Input Reformulations for Translation
- Emergent inabilities? Inverse scaling over the course of pretraining