147 citations · 215 across the 46 of their papers we have counts for
7 papers · 1 filter
Risk Aware Benchmarking of Large Language Models
Apoorva Nitsure, Youssef Mroueh, Mattia Rigotti +6
We propose a distributional framework for benchmarking socio-technical risks of foundation models with quantified statistical significance. Our approach hinges on a new statistical…
Max-Sliced Mutual Information
Dor Tsur, Ziv Goldfeld, Kristjan Greenewald
Quantifying the dependence between high-dimensional random variables is central to statistical learning and inference. Two classical methods are canonical correlation analysis (CCA…
Identifiability Guarantees for Causal Disentanglement from Soft Interventions
Jiaqi Zhang, Chandler Squires, Kristjan Greenewald +3
Causal disentanglement aims to uncover a representation of data using latent variables that are interrelated through a causal model. Such a representation is identifiable if the la…
High-Dimensional Smoothed Entropy Estimation via Dimensionality Reduction
Kristjan Greenewald, Brian Kingsbury, Yuancheng Yu
We study the problem of overcoming exponential sample complexity in differential entropy estimation under Gaussian convolutions. Specifically, we consider the estimation of the dif…
Post-processing Private Synthetic Data for Improving Utility on Selected Measures
Hao Wang, Shivchander Sudalairaj, John Henning +2
Existing private synthetic data generation algorithms are agnostic to downstream tasks. However, end users may have specific requirements that the synthetic data must satisfy. Fail…
Finite sample rates of convergence for the Bigraphical and Tensor graphical Lasso estimators
Shuheng Zhou, Kristjan Greenewald
Many modern datasets exhibit dependencies among observations as well as variables. A decade ago, Kalaitzis et. al. (2013) proposed the Bigraphical Lasso, an estimator for precision…