On false discovery rate thresholding for classification under sparsity
arXiv:1106.6147 · doi:10.1214/12-AOS1042
Abstract
We study the properties of false discovery rate (FDR) thresholding, viewed as a classification procedure. The "0"-class (null) is assumed to have a known density while the "1"-class (alternative) is obtained from the "0"-class either by translation or by scaling. Furthermore, the "1"-class is assumed to have a small number of elements w.r.t. the "0"-class (sparsity). We focus on densities of the Subbotin family, including Gaussian and Laplace models. Nonasymptotic oracle inequalities are derived for the excess risk of FDR thresholding. These inequalities lead to explicit rates of convergence of the excess risk to zero, as the number m of items to be classified tends to infinity and in a regime where the power of the Bayes rule is away from 0 and 1. Moreover, these theoretical investigations suggest an explicit choice for the target level of FDR thresholding, as a function of m. Our oracle inequalities show theoretically that the resulting FDR thresholding adapts to the unknown sparsity regime contained in the data. This property is illustrated with numerical experiments.
References in corpus (9)
- On the Benjamini--Hochberg method
- Microarrays, Empirical Bayes and the Two-Groups Model
- An adaptive step-down procedure with proven FDR control under independence
- On the false discovery rate and an asymptotically optimal rejection curve
- A comparison of the Benjamini-Hochberg procedure with some Bayesian rules for multiple testing
- Asymptotic minimaxity of False Discovery Rate thresholding for sparse exponential data
- Exact calculations for false discovery proportion with application to least favorable configurations
- On the performance of FDR control: Constraints and a partial solution
- Type I error rate control for testing many hypotheses: a survey with proofs