Test Sample Accuracy Scales with Training Sample Density in Neural Networks
arXiv:2106.08365
Abstract
Intuitively, one would expect accuracy of a trained neural network's prediction on test samples to correlate with how densely the samples are surrounded by seen training samples in representation space. We find that a bound on empirical training error smoothed across linear activation regions scales inversely with training sample density in representation space. Empirically, we verify this bound is a strong predictor of the inaccuracy of the network's prediction on test samples. For unseen test sets, including those with out-of-distribution samples, ranking test samples by their local region's error bound and discarding samples with the highest bounds raises prediction accuracy by up to 20% in absolute terms for image classification datasets, on average over thresholds.
CoLLAs 2022 oral
References in corpus (5)
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
- Classification-Based Anomaly Detection for General Data
- SSD: A Unified Framework for Self-Supervised Outlier Detection
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution Generalization