Making Use of NXt to Nothing: The Effect of Class Imbalances on DGA Detection Classifiers
arXiv:2007.00300 · doi:10.1145/3407023.3409190
Abstract
Numerous machine learning classifiers have been proposed for binary classification of domain names as either benign or malicious, and even for multiclass classification to identify the domain generation algorithm (DGA) that generated a specific domain name. Both classification tasks have to deal with the class imbalance problem of strongly varying amounts of training samples per DGA. Currently, it is unclear whether the inclusion of DGAs for which only a few samples are known to the training sets is beneficial or harmful to the overall performance of the classifiers. In this paper, we perform a comprehensive analysis of various contextless DGA classifiers, which reveals the high value of a few training samples per class for both classification tasks. We demonstrate that the classifiers are able to detect various DGAs with high probability by including the underrepresented classes which were previously hardly recognizable. Simultaneously, we show that the classifiers' detection capabilities of well represented classes do not decrease.
Accepted at The 15th International Conference on Availability, Reliability and Security (ARES 2020)
References in corpus (3)
Cited by in corpus (5)
- First Step Towards EXPLAINable DGA Multiclass Classification
- False Sense of Security: Leveraging XAI to Analyze the Reasoning and True Performance of Context-less DGA Classifiers
- The More, the Better? A Study on Collaborative Machine Learning for DGA Detection
- Towards Robust Domain Generation Algorithm Classification
- Detecting Unknown DGAs without Context Information