Multiple testing with the structure adaptive Benjamini-Hochberg algorithm
arXiv:1606.07926
Abstract
In multiple testing problems, where a large number of hypotheses are tested simultaneously, false discovery rate (FDR) control can be achieved with the well-known Benjamini-Hochberg procedure, which adapts to the amount of signal present in the data. Many modifications of this procedure have been proposed to improve power in scenarios where the hypotheses are organized into groups or into a hierarchy, as well as other structured settings. Here we introduce SABHA, the "structure-adaptive Benjamini-Hochberg algorithm", as a generalization of these adaptive testing methods. SABHA incorporates prior information about any pre-determined type of structure in the pattern of locations of the signals and nulls within the list of hypotheses, to reweight the p-values in a data-adaptive way. This raises the power by making more discoveries in regions where signals appear to be more common. Our main theoretical result proves that SABHA controls FDR at a level that is at most slightly higher than the target FDR level, as long as the adaptive weights are constrained sufficiently so as not to overfit too much to the data-interestingly, the excess FDR can be related to the Rademacher complexity or Gaussian width of the class from which we choose our data-adaptive weights. We apply this general framework to various structured settings, including ordered, grouped, and low total variation structures, and get the bounds on FDR for each specific setting. We also examine the empirical performance of SABHA on fMRI activity data and on gene/drug response data, as well as on simulated data.
References in corpus (3)
Cited by in corpus (8)
- Online control of the false discovery rate with decaying memory
- AdaPT: An interactive procedure for multiple testing with side information
- NeuralFDR: Learning Discovery Thresholds from Hypothesis Features
- BONuS: Multiple multivariate testing with a data-adaptivetest statistic
- Contextual Online False Discovery Rate Control
- Analysis of error control in large scale two-stage multiple hypothesis testing
- Two-component Mixture Model in the Presence of Covariates
- The Generic Holdout: Preventing False-Discoveries in Adaptive Data Science