5 papers
Optimal Dimension-Free Sampling for Regularized Classification
Meysam Alishahi, Alexander Munteanu, Simon Omlor +1
We prove optimal sampling bounds achieving -relative error for a broad class of Lipschitz continuous classification loss functions under various regularization t…
Optimizing Computational-Statistical Runtime for Wasserstein Distance Estimation
Peter Matthew Jacobs, Jeff M. Phillips
Squared Wasserstein distance is a frequently used tool to measure discrepancy between probability distributions. This distance is typically computed between empirical measures of s…
TabKDE: Simple and Scalable Tabular Data Generation with Kernel Density Estimates
Meysam Alishahi, Yan Zheng, Junpeng Wang +2
Tabular data generation considers a large table with multiple columns -- each column comprised of numerical, categorical, or sometimes ordinal values. The goal is to produce new ro…
Convolutional Maximum Mean Discrepancy for Inference in Noisy Data
Ritwik Vashistha, Jeff M. Phillips, Abhra Sarkar +1
Modern data analyses frequently encounter settings where samples of variables are contaminated by measurement error. Ignoring measurement noise can substantially degrade statistica…
Hardness of High-Dimensional Linear Classification
Alexander Munteanu, Simon Omlor, Jeff M. Phillips
We establish new exponential in dimension lower bounds for the Maximum Halfspace Discrepancy problem, which models linear classification. Both are fundamental problems in computati…