A systematic comparison of supervised classifiers
arXiv:1311.0202 · doi:10.1371/journal.pone.0094137
Abstract
Pattern recognition techniques have been employed in a myriad of industrial, medical, commercial and academic applications. To tackle such a diversity of data, many techniques have been devised. However, despite the long tradition of pattern recognition research, there is no technique that yields the best classification in all scenarios. Therefore, the consideration of as many as possible techniques presents itself as an fundamental practice in applications aiming at high accuracy. Typical works comparing methods either emphasize the performance of a given algorithm in validation tests or systematically compare various algorithms, assuming that the practical use of these methods is done by experts. In many occasions, however, researchers have to deal with their practical classification tasks without an in-depth knowledge about the underlying mechanisms behind parameters. Actually, the adequate choice of classifiers and parameters alike in such practical circumstances constitutes a long-standing problem and is the subject of the current paper. We carried out a study on the performance of nine well-known classifiers implemented by the Weka framework and compared the dependence of the accuracy with their configuration parameter configurations. The analysis of performance with default parameters revealed that the k-nearest neighbors method exceeds by a large margin the other methods when high dimensional datasets are considered. When other configuration of parameters were allowed, we found that it is possible to improve the quality of SVM in more than 20% even if parameters are set randomly. Taken together, the investigation conducted in this paper suggests that, apart from the SVM implementation, Weka's default configuration of parameters provides an performance close the one achieved with the optimal configuration.
References in corpus (1)
Cited by in corpus (23)
- Text authorship identified using the dynamics of word co-occurrence networks
- A complex network approach to stylometry
- Word sense disambiguation: a complex network approach
- Probing the topological properties of complex networks modeling short written texts
- Classifying informative and imaginative prose using complex networks
- Comparing the writing style of real and artificial papers
- Concentric network symmetry grasps authors' styles in word adjacency networks
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Authorship Attribution Based on Life-Like Network Automata
- Paragraph-based complex networks: application to document classification and authenticity verification
- On the role of words in the network structure of texts: application to authorship attribution
- Topological-collaborative approach for disambiguating authors' names in collaborative networks
- Authorship attribution via network motifs identification
- Physics-integrated machine learning: embedding a neural network in the Navier-Stokes equations. Part I
- Physics-integrated machine learning: embedding a neural network in the Navier-Stokes equations. Part II
- Network analysis of named entity co-occurrences in written texts
- On predicting research grants productivity
- Analyzing the relationship between text features and research proposal productivity
- A pattern recognition approach for distinguishing between prose and poetry
- Improving LBP and its variants using anisotropic diffusion
- Characterizing Complementary Bipolar Junction Transistors by Early Modelling, Image Analysis, and Pattern Recognition
- A comparative analysis of local network similarity measurements: application to author citation networks
- Human activity recognition from skeleton poses