3 papers
math.ST2025
Adversarially robust clustering with optimality guarantees
Soham Jana, Kun Yang, Sanjeev Kulkarni
We consider the problem of clustering data points coming from sub-Gaussian mixtures. Existing methods that provably achieve the optimal mislabeling error, such as the Lloyd algorit…
stat.ML2025
Factor Informed Double Deep Learning For Average Treatment Effect Estimation
Jianqing Fan, Soham Jana, Sanjeev Kulkarni +1
We investigate the problem of estimating the average treatment effect (ATE) under a very general setup where the covariates can be high-dimensional, highly correlated, and can have…
math.ST2024
A provable initialization and robust clustering method for general mixture models
Soham Jana, Jianqing Fan, Sanjeev Kulkarni
Clustering is a fundamental tool in statistical machine learning in the presence of heterogeneous data. Most recent results focus primarily on optimal mislabeling guarantees when d…