Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
arXiv:2106.10328
Abstract
Language models can generate harmful and biased outputs and exhibit undesirable behavior according to a given cultural context. We propose a Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets, an iterative process to significantly change model behavior by crafting and fine-tuning on a dataset that reflects a predetermined set of target values. We evaluate our process using three metrics: quantitative metrics with human evaluations that score output adherence to a target value, toxicity scoring on outputs; and qualitative metrics analyzing the most common word associated with a given social category. Through each iteration, we add additional training dataset examples based on observed shortcomings from evaluations. PALMS performs significantly better on all metrics compared to baseline and control models for a broad range of GPT-3 language model sizes without compromising capability integrity. We find that the effectiveness of PALMS increases with model size. We show that significantly adjusting language model behavior is feasible with a small, hand-curated dataset.
Both authors contributed equally. Accepted at NeurIPS 2021
References in corpus (7)
- Language Models are Few-Shot Learners
- Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence
- Data and its (dis)contents: A survey of dataset development and use in machine learning research
- Alignment of Language Agents
- Extending the Machine Learning Abstraction Boundary: A Complex Systems Approach to Incorporate Societal Context
- Detoxifying Language Models Risks Marginalizing Minority Voices
- Towards Debiasing Sentence Representations
Cited by in corpus (10)
- Collective Constitutional AI: Aligning a Language Model with Public Input
- Finetuned Language Models Are Zero-Shot Learners
- The Future of Intelligent Healthcare: A Systematic Analysis and Discussion on the Integration and Impact of Robots Using Large Language Models for Healthcare
- A General Language Assistant as a Laboratory for Alignment
- Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
- FoRAG: Factuality-optimized Retrieval Augmented Generation for Web-enhanced Long-form Question Answering
- Truthful AI: Developing and governing AI that does not lie
- ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
- SynthBio: A Case Study in Human-AI Collaborative Curation of Text Datasets
- Mitigating harm in language models with conditional-likelihood filtration