A modern maximum-likelihood theory for high-dimensional logistic regression
arXiv:1803.06964 · doi:10.1073/pnas.1810420116
Abstract
Every student in statistics or data science learns early on that when the sample size largely exceeds the number of variables, fitting a logistic model produces estimates that are approximately unbiased. Every student also learns that there are formulas to predict the variability of these estimates which are used for the purpose of statistical inference; for instance, to produce p-values for testing the significance of regression coefficients. Although these formulas come from large sample asymptotics, we are often told that we are on reasonably safe grounds when is large in such a way that or . This paper shows that this is far from the case, and consequently, inferences routinely produced by common software packages are often unreliable. Consider a logistic model with independent features in which and become increasingly large in a fixed ratio. Then we show that (1) the MLE is biased, (2) the variability of the MLE is far greater than classically predicted, and (3) the commonly used likelihood-ratio test (LRT) is not distributed as a chi-square. The bias of the MLE is extremely problematic as it yields completely wrong predictions for the probability of a case based on observed values of the covariates. We develop a new theory, which asymptotically predicts (1) the bias of the MLE, (2) the variability of the MLE, and (3) the distribution of the LRT. We empirically also demonstrate that these predictions are extremely accurate in finite samples. Further, an appealing feature is that these novel predictions depend on the unknown sequence of regression coefficients only through a single scalar, the overall strength of the signal. This suggests very concrete procedures to adjust inference; we describe one such procedure learning a single parameter from data and producing accurate inference
29 pages, 14 figures, 4 tables
References in corpus (1)
Cited by in corpus (36)
- The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime
- Global and Simultaneous Hypothesis Testing for High-Dimensional Logistic Regression Models
- The Asymptotic Distribution of the MLE in High-dimensional Logistic Models: Arbitrary Covariance
- The Impact of Regularization on High-dimensional Logistic Regression
- A Precise High-Dimensional Asymptotic Theory for Boosting and Minimum--Norm Interpolated Classifiers
- Finite-sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime
- A Model of Double Descent for High-dimensional Binary Linear Classification
- Hierarchical inference for genome-wide association studies: a view on methodology with software
- Don't Just Blame Over-parametrization for Over-confidence: Theoretical Analysis of Calibration in Binary Classification
- Improving the Stability of the Knockoff Procedure: Multiple Simultaneous Knockoffs and Entropy Maximization
- Learning from Binary Multiway Data: Probabilistic Tensor Decomposition and its Statistical Optimality
- Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation
- Online stochastic gradient descent on non-convex losses from high-dimensional inference
- Analysis of overfitting in the regularized Cox model
- Label-Imbalanced and Group-Sensitive Classification under Overparameterization
- High Dimensional Classification via Regularized and Unregularized Empirical Risk Minimization: Precise Error and Optimal Loss
- Probabilistic learning inference of boundary value problem with uncertainties based on Kullback-Leibler divergence under implicit constraints
- Regularization in High-Dimensional Regression and Classification via Random Matrix Theory
- Strong replica symmetry for high-dimensional disordered log-concave Gibbs measures
- Tractability from overparametrization: The example of the negative perceptron
- The role of regularization in classification of high-dimensional noisy Gaussian mixture
- Theoretical characterization of uncertainty in high-dimensional linear classification
- Consistency of semi-supervised learning algorithms on graphs: Probit and one-hot methods
- Penalization-induced shrinking without rotation in high dimensional GLM regression: a cavity analysis
- Tensor denoising and completion based on ordinal observations
- Exact high-dimensional asymptotics for Support Vector Machine
- SLOE: A Faster Method for Statistical Inference in High-Dimensional Logistic Regression
- Directional testing for high-dimensional multivariate normal distributions
- Solvable Model for Inheriting the Regularization through Knowledge Distillation
- On the Properties of Simulation-based Estimators in High Dimensions
- StarTrek: Combinatorial Variable Selection with False Discovery Rate Control
- High Dimensional Logistic Regression Under Network Dependence
- Towards Designing Optimal Sensing Matrices for Generalized Linear Inverse Problems
- Phase Transition Unbiased Estimation in High Dimensional Settings
- De-biased Lasso for Generalized Linear Models with A Diverging Number of Covariates
- The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models