Measuring support for a hypothesis about a random parameter without estimating its unknown prior
arXiv:1101.0305 · doi:10.1111/insr.12008
Abstract
For frequentist settings in which parameter randomness represents variability rather than uncertainty, the ideal measure of the support for one hypothesis over another is the difference in the posterior and prior log odds. For situations in which the prior distribution cannot be accurately estimated, that ideal support may be replaced by another measure of support, which may be any predictor of the ideal support that, on a per-observation basis, is asymptotically unbiased. Two qualifying measures of support are defined. The first is minimax optimal with respect to the population and is equivalent to a particular Bayes factor. The second is worst-sample minimax optimal and is equivalent to the normalized maximum likelihood. It has been extended by likelihood weights for compatibility with more general models. One such model is that of two independent normal samples, the standard setting for gene expression microarray data analysis. Applying that model to proteomics data indicates that support computed from data for a single protein can closely approximate the estimated difference in posterior and prior odds that would be available with the data for 20 proteins. This suggests the applicability of random-parameter models to other situations in which the parameter distribution cannot be reliably estimated.
Errors in the first version were corrected, and the methodology is now applied to more interesting data
References in corpus (9)
- Simultaneous inference: When should hypothesis testing problems be combined?
- The Future of Indirect Evidence
- Statistical inference optimized with respect to the observed sample for single or multiple comparisons
- Empirical Bayes estimation of posterior probabilities of enrichment
- A Law of Likelihood for Composite Hypotheses
- Resolving conflicts between statistical methods by probability combination: Application to empirical Bayes analyses of genomic data
- Large-scale interval and point estimates from an empirical Bayes extension of confidence posteriors
- Controlling the degree of caution in statistical inference with the Bayesian and frequentist approaches as opposite extremes
- Simple estimators of false discovery rates given as few as one or two p-values without strong parametric assumptions