MLDemon: Deployment Monitoring for Machine Learning Systems
arXiv:2104.13621
Abstract
Post-deployment monitoring of ML systems is critical for ensuring reliability, especially as new user inputs can differ from the training distribution. Here we propose a novel approach, MLDemon, for ML DEployment MONitoring. MLDemon integrates both unlabeled data and a small amount of on-demand labels to produce a real-time estimate of the ML model's current performance on a given data stream. Subject to budget constraints, MLDemon decides when to acquire additional, potentially costly, expert supervised labels to verify the model. On temporal datasets with diverse distribution drifts and models, MLDemon outperforms existing approaches. Moreover, we provide theoretical analysis to show that MLDemon is minimax rate optimal for a broad class of distribution drifts.
Accepted to AISTATS 2022. Significant changes to algorithm, theory, and experiments since previous versions
References in corpus (8)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Do ImageNet Classifiers Generalize to ImageNet?
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- On Learning Invariant Representation for Domain Adaptation
- Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL
- The Effect of Natural Distribution Shift on Question Answering Models
- FrugalML: How to Use ML Prediction APIs More Accurately and Cheaply
- Feature Shift Detection: Localizing Which Features Have Shifted via Conditional Distribution Tests