SoK: Machine Learning Governance
arXiv:2109.10870
Abstract
The application of machine learning (ML) in computer systems introduces not only many benefits but also risks to society. In this paper, we develop the concept of ML governance to balance such benefits and risks, with the aim of achieving responsible applications of ML. Our approach first systematizes research towards ascertaining ownership of data and models, thus fostering a notion of identity specific to ML systems. Building on this foundation, we use identities to hold principals accountable for failures of ML systems through both attribution and auditing. To increase trust in ML systems, we then survey techniques for developing assurance, i.e., confidence that the system meets its security requirements and does not exhibit certain known failures. This leads us to highlight the need for techniques that allow a model owner to manage the life cycle of their system, e.g., to patch or retire their ML system. Put altogether, our systematization of knowledge standardizes the interactions between principals involved in the deployment of ML throughout its life cycle. We highlight opportunities for future work, e.g., to formalize the resulting game between ML principals.
References in corpus (28)
- Language Models are Few-Shot Learners
- Equality of Opportunity in Supervised Learning
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Certified Adversarial Robustness via Randomized Smoothing
- Exponentially Weighted Moving Average Charts for Detecting Concept Drift
- Deep Anomaly Detection with Outlier Exposure
- Practical Secure Aggregation for Federated Learning on User-Held Data
- On the Convergence and Robustness of Adversarial Training
- Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition
- Transferable Clean-Label Poisoning Attacks on Deep Neural Nets
- Recent Advances in Algorithmic High-Dimensional Robust Statistics
- Auditing Differentially Private Machine Learning: How Private is Private SGD?
- On the Effectiveness of Mitigating Data Poisoning Attacks with Gradient Shaping
- Detecting GAN-generated Imagery using Color Cues
- Differentially Private Learning Needs Better Features (or Much More Data)
- Adaptive Machine Unlearning
- Encode, Shuffle, Analyze Privacy Revisited: Formalizations and Empirical Evaluation
- Fairer and more accurate, but for whom?
- Dataset Inference: Ownership Resolution in Machine Learning
- Detecting Out-of-Distribution Examples with In-distribution Examples and Gram Matrices
- MT-Adapted Datasheets for Datasets: Template and Repository
- Analysis of Confident-Classifiers for Out-of-distribution Detection
- An Ethical Highlighter for People-Centric Dataset Creation
- Secure Medical Image Analysis with CrypTFlow
- On the Privacy Risks of Algorithmic Fairness
- Proof-of-Learning: Definitions and Practice
- Do GANs leave artificial fingerprints?
- On the Exploitability of Audio Machine Learning Pipelines to Surreptitious Adversarial Examples