Statistical process monitoring of artificial neural networks
arXiv:2209.07436 · doi:10.1080/00401706.2023.2239886
Abstract
The rapid advancement of models based on artificial intelligence demands innovative monitoring techniques which can operate in real time with low computational costs. In machine learning, especially if we consider artificial neural networks (ANNs), the models are often trained in a supervised manner. Consequently, the learned relationship between the input and the output must remain valid during the model's deployment. If this stationarity assumption holds, we can conclude that the ANN provides accurate predictions. Otherwise, the retraining or rebuilding of the model is required. We propose considering the latent feature representation of the data (called "embedding") generated by the ANN to determine the time when the data stream starts being nonstationary. In particular, we monitor embeddings by applying multivariate control charts based on the data depth calculation and normalized ranks. The performance of the introduced method is compared with benchmark approaches for various ANN architectures and different underlying data formats.
References in corpus (10)
- Learning under Concept Drift: A Review
- Generalized Out-of-Distribution Detection: A Survey
- Exploring the Limits of Out-of-Distribution Detection
- Addressing Failure Prediction by Learning Model Confidence
- Robust Machine Learning Applied to Astronomical Datasets I: Star-Galaxy Classification of the SDSS DR3 Using Decision Trees
- OpenOOD: Benchmarking Generalized Out-of-Distribution Detection
- Understanding Softmax Confidence and Uncertainty
- Is Out-of-Distribution Detection Learnable?
- Data Twinning
- Anomaly detection using data depth: multivariate case