Concentration inequalities and asymptotic results for ratio type empirical processes
arXiv:math/0606788 · doi:10.1214/009117906000000070
Abstract
Let be a class of measurable functions on a measurable space with values in and let \[P_n=n^{-1}\sum_{i=1}^nδ_{X_i}\] be the empirical measure based on an i.i.d. sample from a probability distribution on . We study the behavior of suprema of the following type: \[\sup_{r_n<σ_Pf\leq δ_n}\frac{|P_nf-Pf|}{ϕ(σ_Pf)},\] where and is a continuous, strictly increasing function with . Using Talagrand's concentration inequality for empirical processes, we establish concentration inequalities for such suprema and use them to derive several results about their asymptotic behavior, expressing the conditions in terms of expectations of localized suprema of empirical processes. We also prove new bounds for expected values of sup-norms of empirical processes in terms of the largest and the norm of the envelope of the function class, which are especially suited for estimating localized suprema. With this technique, we extend to function classes most of the known results on ratio type suprema of empirical processes, including some of Alexander's results for VC classes of sets. We also consider applications of these results to several important problems in nonparametric statistics and in learning theory (including general excess risk bounds in empirical risk minimization and their versions for -regression and classification and ratio type bounds for margin distributions in classification).
Published at http://dx.doi.org/10.1214/009117906000000070 in the Annals of Probability (http://www.imstat.org/aop/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (2)
Cited by in corpus (48)
- 2004 IMS Medallion Lecture: Local Rademacher complexities and oracle inequalities in risk minimization
- Nonparametric regression using deep neural networks with ReLU activation function
- Moving Beyond Sub-Gaussianity in High-Dimensional Statistics: Applications in Covariance Estimation and Linear Regression
- On local -statistic processes and the estimation of densities of functions of several sample variables
- The Optimal Sample Complexity of PAC Learning
- Activized Learning: Transforming Passive to Active with Improved Label Complexity
- A new method for estimation and model selection: -estimation
- Rates of convergence in active learning
- Uniform in bandwidth consistency of local polynomial regression function estimators
- On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces
- Adaptive estimation of a distribution function and its density in sup-norm loss by wavelet and spline projections
- Minimax Confidence Intervals for the Sliced Wasserstein Distance
- On multivariate quantiles under partial orders
- Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
- Surrogate Losses in Passive and Active Learning
- Uniform bounds for norms of sums of independent random functions
- On Equivalence of Martingale Tail Bounds and Deterministic Regret Inequalities
- Nonparametric Heterogeneity Testing For Massive Data
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural network
- Exponential Savings in Agnostic Active Learning through Abstention
- Robustness of shape-restricted regression estimators: an envelope perspective
- Optimal learning via local entropies and sample compression
- Uniform central limit theorems for the Grenander estimator
- Convergence rates of least squares regression estimators with heavy-tailed errors
- Spectral Pruning: Compressing Deep Neural Networks via Spectral Analysis and its Generalization Error
- Fast learning rate of deep learning via a kernel perspective
- Towards Optimal Problem Dependent Generalization Error Bounds in Statistical Learning Theory
- A Compression Technique for Analyzing Disagreement-Based Active Learning
- Complex sampling designs: uniform limit theorems and applications
- Nonparametric Instrumental Variables Estimation Under Misspecification
- Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games
- Global testing against sparse alternatives in time-frequency analysis
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
- Upper functions for positive random functionals
- When are epsilon-nets small?
- Approximating unit balls via random sampling
- Policy Transforms and Learning Optimal Policies
- Generalization Error Estimates of Machine Learning Methods for Solving High Dimensional Schrödinger Eigenvalue Problems
- Two-level monotonic multistage recommender systems
- Upper functions for -norm of gaussian random fields
- Some New Asymptotic Theory for Least Squares Series: Pointwise and Uniform Results
- Adaptive confidence sets in L^2
- A local maximal inequality under uniform entropy
- On the Hermite spline conjecture and its connection to k-monotone densities
- Efficient Simulation-Based Minimum Distance Estimation and Indirect Inference
- Cox process functional learning
- Multiplier U-processes: sharp bounds and applications
- On the uniform convergence of empirical norms and inner products, with application to causal inference