Auditing large language models: a three-layered approach
arXiv:2302.08500 · doi:10.1007/s43681-023-00289-2
Abstract
Large language models (LLMs) represent a major advance in artificial intelligence (AI) research. However, the widespread use of LLMs is also coupled with significant ethical and social challenges. Previous research has pointed towards auditing as a promising governance mechanism to help ensure that AI systems are designed and deployed in ways that are ethical, legal, and technically robust. However, existing auditing procedures fail to address the governance challenges posed by LLMs, which display emergent capabilities and are adaptable to a wide range of downstream tasks. In this article, we address that gap by outlining a novel blueprint for how to audit LLMs. Specifically, we propose a three-layered approach, whereby governance audits (of technology providers that design and disseminate LLMs), model audits (of LLMs after pre-training but prior to their release), and application audits (of applications based on LLMs) complement and inform each other. We show how audits, when conducted in a structured and coordinated manner on all three levels, can be a feasible and effective mechanism for identifying and managing some of the ethical and social risks posed by LLMs. However, it is important to remain realistic about what auditing can reasonably be expected to achieve. Therefore, we discuss the limitations not only of our three-layered approach but also of the prospect of auditing LLMs at all. Ultimately, this article seeks to expand the methodological toolkit available to technology providers and policymakers who wish to analyse and evaluate LLMs from technical, ethical, and legal perspectives.
22 pages, 2 figures. AI Ethics (2023)
References in corpus (24)
- Learning Transferable Visual Models From Natural Language Supervision
- Training language models to follow instructions with human feedback
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Scaling Laws for Neural Language Models
- Society-in-the-Loop: Programming the Algorithmic Social Contract
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- The Fallacy of AI Functionality
- Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning
- Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem
- Conformity Assessments and Post-market Monitoring: A Guide to the Role of Auditing in the Proposed European AI Regulation
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Operationalising AI governance through ethics-based auditing: An industry case study
- Interactive Model Cards: A Human-Centered Approach to Model Documentation
- The US Algorithmic Accountability Act of 2022 vs. The EU Artificial Intelligence Act: What can they learn from each other?
- AI Certification: Advancing Ethical Practice by Reducing Information Asymmetries
- Ethics-Based Auditing of Automated Decision-Making Systems: Intervention Points and Policy Implications
- Filling gaps in trustworthy development of AI
- Discovering Language Model Behaviors with Model-Written Evaluations
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- Epistemic values in feature importance methods: Lessons from feminist epistemology
- Challenges and Best Practices in Corporate AI Governance:Lessons from the Biopharmaceutical Industry
- Red Teaming Language Models with Language Models
- Contrastive Adapters for Foundation Model Group Robustness
- Language Models are Good Translators
Cited by in corpus (16)
- Auditing of AI: Legal, Ethical and Technical Approaches
- Black-Box Access is Insufficient for Rigorous AI Audits
- LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
- Exploring the Roles of Large Language Models in Reshaping Transportation Systems: A Survey, Framework, and Roadmap
- Three lines of defense against risks from AI
- Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
- A five-layer framework for AI governance: integrating regulation, standards, and certification
- Shaping New Norms for AI
- Frontier AI developers need an internal audit function
- Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
- Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
- Can ChatGPT Perform Reasoning Using the IRAC Method in Analyzing Legal Scenarios Like a Lawyer?
- Exploring ChatGPT's Capabilities, Stability, Potential and Risks in Conducting Psychological Counseling through Simulations in School Counseling
- Mapping Human Anti-collusion Mechanisms to Multi-agent AI Systems
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study
- Towards AI epidemiology: a measurement standardisation framework for prospective risk detection