ECOD: Unsupervised Outlier Detection Using Empirical Cumulative Distribution Functions
arXiv:2201.00382 · doi:10.1109/TKDE.2022.3159580
Abstract
Outlier detection refers to the identification of data points that deviate from a general data distribution. Existing unsupervised approaches often suffer from high computational cost, complex hyperparameter tuning, and limited interpretability, especially when working with large, high-dimensional datasets. To address these issues, we present a simple yet effective algorithm called ECOD (Empirical-Cumulative-distribution-based Outlier Detection), which is inspired by the fact that outliers are often the "rare events" that appear in the tails of a distribution. In a nutshell, ECOD first estimates the underlying distribution of the input data in a nonparametric fashion by computing the empirical cumulative distribution per dimension of the data. ECOD then uses these empirical distributions to estimate tail probabilities per dimension for each data point. Finally, ECOD computes an outlier score of each data point by aggregating estimated tail probabilities across dimensions. Our contributions are as follows: (1) we propose a novel outlier detection method called ECOD, which is both parameter-free and easy to interpret; (2) we perform extensive experiments on 30 benchmark datasets, where we find that ECOD outperforms 11 state-of-the-art baselines in terms of accuracy, efficiency, and scalability; and (3) we release an easy-to-use and scalable (with distributed support) Python implementation for accessibility and reproducibility.
Accepted to IEEE Transactions on Knowledge and Data Engineering (TKDE) with fixed data statistics. Zheng Li and Yue Zhao contributed equally. Code is available in PyOD library at https://github.com/yzhao062/pyod
References in corpus (2)
Cited by in corpus (12)
- Deep Isolation Forest for Anomaly Detection
- Consistency-guided semi-supervised outlier detection in heterogeneous data using fuzzy rough sets
- Outlier detection in mixed-attribute data: a semi-supervised approach with fuzzy approximations and relative entropy
- Quadratic Neuron-empowered Heterogeneous Autoencoder for Unsupervised Anomaly Detection
- Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data
- Robust Outlier Detection Method Based on Local Entropy and Global Density
- Label-Informed Outlier Detection Based on Granule Density
- Neural Collaborative Filtering to Detect Anomalies in Human Semantic Trajectories
- Unsupervised Surrogate Anomaly Detection
- Flexible Simulation Based Inference for Galaxy Photometric Fitting with Synthesizer
- Low-count Time Series Anomaly Detection
- Random Similarity Isolation Forests