A Framework for Cluster and Classifier Evaluation in the Absence of Reference Labels
arXiv:2109.11126 · doi:10.1145/3474369.3486867
Abstract
In some problem spaces, the high cost of obtaining ground truth labels necessitates use of lower quality reference datasets. It is difficult to benchmark model performance using these datasets, as evaluation results may be biased. We propose a supplement to using reference labels, which we call an approximate ground truth refinement (AGTR). Using an AGTR, we prove that bounds on specific metrics used to evaluate clustering algorithms and multi-class classifiers can be computed without reference labels. We also introduce a procedure that uses an AGTR to identify inaccurate evaluation results produced from datasets of dubious quality. Creating an AGTR requires domain knowledge, and malware family classification is a task with robust domain knowledge approaches that support the construction of an AGTR. We demonstrate our AGTR evaluation framework by applying it to a popular malware labeling tool to diagnose over-fitting in prior testing and evaluate changes whose impact could not be meaningfully quantified under previous data.
to appear in Proceedings of the 14th ACM Workshop on Artificial Intelligence and Security
References in corpus (11)
- Microsoft Malware Classification Challenge
- Realistic Evaluation of Deep Semi-Supervised Learning Algorithms
- EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models
- Do We Train on Test Data? Purging CIFAR of Near-Duplicates
- A Step Toward Quantifying Independently Reproducible Machine Learning Research
- Accounting for Variance in Machine Learning Benchmarks
- A Survey of Machine Learning Methods and Challenges for Windows Malware Classification
- Assessing Generalization of SGD via Disagreement
- Leveraging Uncertainty for Improved Static Malware Detection Under Extreme False Positive Constraints
- Identifying Statistical Bias in Dataset Replication
- Are Labels Always Necessary for Classifier Accuracy Evaluation?
Cited by in corpus (5)
- The Cross-evaluation of Machine Learning-based Network Intrusion Detection Systems
- Logical Assessment Formula and Its Principles for Evaluations with Inaccurate Ground-Truth Labels
- Understanding the Process of Data Labeling in Cybersecurity
- Validation of the Practicability of Logical Assessment Formula for Evaluations with Inaccurate Ground-Truth Labels: An Application Study on Tumour Segmentation for Breast Cancer
- Zipf-Gramming: Scaling Byte N-Grams Up to Production Sized Malware Corpora