Selection of variables for cluster analysis and classification rules
arXiv:math/0610757 · doi:10.1198/016214508000000544
Abstract
In this paper we introduce two procedures for variable selection in cluster analysis and classification rules. One is mainly oriented to detect the noisy non-informative variables, while the other deals also with multicolinearity. A forward-backward algorithm is also proposed to make feasible these procedures in large data sets. A small simulation is performed and some real data examples are analyzed.
28 pages, 7 figures