An Empirical Evaluation of Similarity Measures for Time Series Classification
arXiv:1401.3973 · doi:10.1016/j.knosys.2014.04.035
Abstract
Time series are ubiquitous, and a measure to assess their similarity is a core part of many computational systems. In particular, the similarity measure is the most essential ingredient of time series clustering and classification systems. Because of this importance, countless approaches to estimate time series similarity have been proposed. However, there is a lack of comparative studies using empirical, rigorous, quantitative, and large-scale assessment strategies. In this article, we provide an extensive evaluation of similarity measures for time series classification following the aforementioned principles. We consider 7 different measures coming from alternative measure `families', and 45 publicly-available time series data sets coming from a wide variety of scientific domains. We focus on out-of-sample classification accuracy, but in-sample accuracies and parameter choices are also discussed. Our work is based on rigorous evaluation methodologies and includes the use of powerful statistical significance tests to derive meaningful conclusions. The obtained results show the equivalence, in terms of accuracy, of a number of measures, but with one single candidate outperforming the rest. Such findings, together with the followed methodology, invite researchers on the field to adopt a more consistent evaluation criteria and a more informed decision regarding the baseline measures to which new developments should be compared.
28 pages, 5 figures, 3 tables
References in corpus (1)
Cited by in corpus (16)
- Clustering Based Feature Learning on Variable Stars
- Synthesis of Realistic ECG using Generative Adversarial Networks
- Particle swarm optimization for time series motif discovery
- Forecasting the abnormal events at well drilling with machine learning
- The Fraction of Broken Waves in Natural Surf Zones
- Sparsification of the Alignment Path Search Space in Dynamic Time Warping
- Automatic Registration and Clustering of Time Series
- Ranking and significance of variable-length similarity-based time series motifs
- Autoencoder Based Iterative Modeling and Multivariate Time-Series Subsequence Clustering Algorithm
- Modeling Financial Time Series using LSTM with Trainable Initial Hidden States
- Application of Machine Learning to accidents detection at directional drilling
- Synthetic Random Environmental Time Series Generation with Similarity Control, Preserving Original Signal's Statistical Characteristics
- Modelling antimicrobial prescriptions in Scotland: A spatio-temporal clustering approach
- Improving state estimation through projection post-processing for activity recognition with application to football
- Dynamic Time Warp Convolutional Networks
- Report: Dynamic Eye Movement Matching and Visualization Tool in Neuro Gesture