An index of effective number of variables for uncertainty and reliability analysis in model selection problems
arXiv:2602.21403 · doi:10.1016/j.sigpro.2024.109735
Abstract
An index of an effective number of variables (ENV) is introduced for model selection in nested models. This is the case, for instance, when we have to decide the order of a polynomial function or the number of bases in a nonlinear regression, choose the number of clusters in a clustering problem, or the number of features in a variable selection application (to name few examples). It is inspired by the idea of the maximum area under the curve (AUC). The interpretation of the ENV index is identical to the effective sample size (ESS) indices concerning a set of samples. The ENV index improves {drawbacks of} the elbow detectors described in the literature and introduces different confidence measures of the proposed solution. These novel measures can be also employed jointly with the use of different information criteria, such as the well-known AIC and BIC, or any other model selection procedures. Comparisons with classical and recent schemes are provided in different experiments involving real datasets. Related Matlab code is given.
References in corpus (8)
- Model Order Selection Based on Information Theoretic Criteria: Design of the Penalty
- On the safe use of prior densities for Bayesian model selection
- Spectral information criterion for automatic elbow detection
- Universal and Automatic Elbow Detection for Learning the Effective Number of Components in Model Selection Problems
- An exhaustive variable selection study for linear models of soundscape emotions: rankings and Gibbs analysis
- Monte-Carlo Sampling Approach to Model Selection: A Primer
- Asymptotic efficiency for Sobol' and Cram{é}r-von Mises indices under two designs of experiments
- A Generalized Variable Importance Metric and Estimator for Black Box Machine Learning Models