Collective Constitutional AI: Aligning a Language Model with Public Input
arXiv:2406.07814 · doi:10.1145/3630106.3658979
Abstract
There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To address this need, we present Collective Constitutional AI (CCAI): a multi-stage process for sourcing and integrating public input into LMs-from identifying a target population to sourcing principles to training and evaluating a model. We demonstrate the real-world practicality of this approach by creating what is, to our knowledge, the first LM fine-tuned with collectively sourced public input and evaluating this model against a baseline model trained with established principles from a LM developer. Our quantitative evaluations demonstrate several benefits of our approach: the CCAI-trained model shows lower bias across nine social dimensions compared to the baseline model, while maintaining equivalent performance on language, math, and helpful-harmless evaluations. Qualitative comparisons of the models suggest that the models differ on the basis of their respective constitutions, e.g., when prompted with contentious topics, the CCAI-trained model tends to generate responses that reframe the matter positively instead of a refusal. These results demonstrate a promising, tractable pathway toward publicly informed development of language models.
References in corpus (6)
- Training language models to follow instructions with human feedback
- Measurement and Fairness
- Jury Learning: Integrating Dissenting Voices into Machine Learning Models
- AI's Regimes of Representation: A Community-centered Study of Text-to-Image Models in South Asia
- Queer In AI: A Case Study in Community-Led Participatory AI
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
Cited by in corpus (10)
- Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
- Beyond Preferences in AI Alignment
- Perceptions of Sentient AI and Other Digital Minds: Evidence from the AI, Morality, and Sentience (AIMS) Survey
- C3AI: Crafting and Evaluating Constitutions for Constitutional AI
- Ontologies in Design: How Imagining a Tree Reveals Possibilites and Assumptions in Large Language Models
- Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
- Case Law Grounding: Using Precedents to Align Decision-Making for Humans and AI
- Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
- AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence
- Statutory AI: Aligning Large Language Models With Legal Norms