Model Selection Principles in Misspecified Models
arXiv:1005.5483 · doi:10.1111/rssb.12023
Abstract
Model selection is of fundamental importance to high dimensional modeling featured in many contemporary applications. Classical principles of model selection include the Kullback-Leibler divergence principle and the Bayesian principle, which lead to the Akaike information criterion and Bayesian information criterion when models are correctly specified. Yet model misspecification is unavoidable when we have no knowledge of the true model or when we have the correct family of distributions but miss some true predictor. In this paper, we propose a family of semi-Bayesian principles for model selection in misspecified models, which combine the strengths of the two well-known principles. We derive asymptotic expansions of the semi-Bayesian principles in misspecified generalized linear models, which give the new semi-Bayesian information criteria (SIC). A specific form of SIC admits a natural decomposition into the negative maximum quasi-log-likelihood, a penalty on model dimensionality, and a penalty on model misspecification directly. Numerical studies demonstrate the advantage of the newly proposed SIC methodology for model selection in both correctly specified and misspecified models.
25 pages, 6 tables
References in corpus (4)
Cited by in corpus (6)
- Scientific discovery in a model-centric framework: Reproducibility, innovation, and epistemic diversity
- Regularization Methods for High-Dimensional Instrumental Variables Regression With an Application to Genetical Genomics
- Model-free Feature Screening and FDR Control with Knockoff Features
- High-dimensional variable selection via low-dimensional adaptive learning
- Discussion: "A significance test for the lasso"
- Selecting fitted models under epistemic uncertainty using a stochastic process on quantile functions