Outlier Detection in High Dimensional Data
arXiv:1909.03681 · doi:10.1142/S0219649220400134
Abstract
High-dimensional data poses unique challenges in outlier detection process. Most of the existing algorithms fail to properly address the issues stemming from a large number of features. In particular, outlier detection algorithms perform poorly on data set of small size with a large number of features. In this paper, we propose a novel outlier detection algorithm based on principal component analysis and kernel density estimation. The proposed method is designed to address the challenges of dealing with high-dimensional data by projecting the original data onto a smaller space and using the innate structure of the data to calculate anomaly scores for each data point. Numerical experiments on synthetic and real-life data show that our method performs well on high-dimensional data. In particular, the proposed method outperforms the benchmark methods as measured by the -score. Our method also produces better-than-average execution times compared to the benchmark methods.
References in corpus (2)
Cited by in corpus (6)
- Kernel density estimation based sampling for imbalanced class distribution
- Gamma distribution-based sampling for imbalanced data
- Stock price forecast with deep learning
- Forecasting with Deep Learning: S&P 500 index
- Identifying outliers in astronomical images with unsupervised machine learning
- Identification of Anomalous E+A Galaxies in GAMA Using an Isolation Forest