11 papers
The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression
Kevin Han Huang, Haoyu Ye, Somak Laha +1
Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance…
Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation
Matthew Esmaili Mallory, Kevin Han Huang, Morgane Austern
Over the last decade, a wave of research has characterized the exact asymptotic risk of many high-dimensional models in the proportional regime. Two foundational results have drive…
Computable Bounds for Strong Approximations with Applications
Haoyu Ye, Morgane Austern
The Komlós$\unicode{x2013}$Major$\unicode{x2013}$Tusnády (KMT) inequality for partial sums is one of the most celebrated results in probability theory. Yet its practical applicat…
Inference on Optimal Policy Values and Other Irregular Functionals via Softmax Smoothing
Justin Whitehouse, Qizhao Chen, Morgane Austern +1
Constructing confidence intervals for the value of an (unknown) optimal treatment policy is a fundamental problem in causal inference. Insight into the optimal policy value can gui…
Graph Attention Network for Node Regression on Random Geometric Graphs with ErdÅs--Rényi contamination
Somak Laha, Suqi Liu, Morgane Austern
Graph attention networks (GATs) are widely used and often appear robust to noise in node covariates and edges, yet rigorous statistical guarantees demonstrating a provable advantag…
Poisson-Process Topic Model for Integrating Knowledge from Pre-trained Language Models
Morgane Austern, Yuanchuan Guo, Zheng Tracy Ke +1
Topic modeling is traditionally applied to word counts without accounting for the context in which words appear. Recent advancements in large language models (LLMs) offer contextua…