papers

Publications (45)

cs.LG2025

Hamiltonian Monte Carlo Inference of Marginalized Linear Mixed-Effects Models

Jinlin Lai, Justin Domke, Daniel Sheldon

Bayesian reasoning in linear mixed-effects models (LMMs) is challenging and often requires advanced sampling techniques like Markov chain Monte Carlo (MCMC). A common approach is t…

cs.DB2024

AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic Data

Ryan McKenna, Brett Mullins, Daniel Sheldon +1

We propose AIM, a new algorithm for differentially private synthetic data generation. AIM is a workload-adaptive algorithm within the paradigm of algorithms that first selects a se…

cs.LG2020

Divide and Couple: Using Monte Carlo Variational Objectives for Posterior Approximation

Justin Domke, Daniel Sheldon

Recent work in variational inference (VI) uses ideas from Monte Carlo estimation to tighten the lower bounds on the log-likelihood that are used as objectives. However, there is no…

cs.LG2023

Kernel Interpolation with Sparse Grids

Mohit Yadav, Daniel Sheldon, Cameron Musco

Structured kernel interpolation (SKI) accelerates Gaussian process (GP) inference by interpolating the kernel covariance function using a dense grid of inducing points, whose corre…

stat.ME2020

Three-quarter Sibling Regression for Denoising Observational Data

Shiv Shankar, Daniel Sheldon, Tao Sun +2

Many ecological studies and conservation policies are based on field observations of species, which can be affected by systematic variability introduced by the observation process.…

cs.CV2020

Detecting and Tracking Communal Bird Roosts in Weather Radar Data

Zezhou Cheng, Saadia Gabriel, Pankaj Bhambhani +4

The US weather radar archive holds detailed information about biological phenomena in the atmosphere over the last 20 years. Communally roosting birds congregate in large numbers a…

cs.CC2013

Hamming Approximation of NP Witnesses

Daniel Sheldon, Neal E. Young

Given a satisfiable 3-SAT formula, how hard is it to find an assignment to the variables that has Hamming distance at most n/2 to a satisfying assignment? More generally, consider…

math.PR2011

First Passage Time of Skew Brownian Motion

Thilanka Appuhamillage, Daniel Sheldon

Nearly fifty years after the introduction of skew Brownian motion by Itô and McKean (1963), the first passage time distribution remains unknown. In this paper, we generalize resul…

cs.LG2018

Importance Weighting and Variational Inference

Justin Domke, Daniel Sheldon

Recent work used importance sampling ideas for better variational bounds on likelihoods. We clarify the applicability of these ideas to pure probabilistic inference, by showing the…

cs.LG2014

Gaussian Approximation of Collective Graphical Models

Li-Ping Liu, Daniel Sheldon, Thomas G. Dietterich

The Collective Graphical Model (CGM) models a population of independent and identically distributed individuals when only collective statistics (i.e., counts of individuals) are ob…

cs.CV2023

DISCount: Counting in Large Image Collections with Detector-Based Importance Sampling

Gustavo Perez, Subhransu Maji, Daniel Sheldon

Many modern applications use computer vision to detect and count objects in massive image collections. However, when the detection task is very difficult or in the presence of doma…

cs.LG2023

U-Statistics for Importance-Weighted Variational Inference

Javier Burroni, Kenta Takatsu, Justin Domke +1

We propose the use of U-statistics to reduce variance for gradient estimation in importance-weighted variational inference. The key observation is that, given a base gradient estim…

cs.LG2017

Differentially Private Learning of Undirected Graphical Models using Collective Graphical Models

Garrett Bernstein, Ryan McKenna, Tao Sun +3

We investigate the problem of learning discrete, undirected graphical models in a differentially private way. We show that the approach of releasing noisy sufficient statistics usi…

cs.CV2019

A Bayesian Perspective on the Deep Image Prior

Zezhou Cheng, Matheus Gadelha, Subhransu Maji +1

The deep image prior was recently introduced as a prior for natural images. It represents images as the output of a convolutional network with random inputs. For "inference", gradi…

cs.LG2026

Private Adaptive Covariance Estimation via Gaussian Graphical Models

Cecilia Ferrando, Miguel Fuentes, Brett Mullins +2

We propose PACE-GGM, a data-adaptive differentially private method for covariance estimation that concentrates its privacy budget on the most informative entries of the empirical c…

stat.ML2020

Normalizing Flows Across Dimensions

Edmond Cunningham, Renos Zabounidis, Abhinav Agrawal +2

Real-world data with underlying structure, such as pictures of faces, are hypothesized to lie on a low-dimensional manifold. This manifold hypothesis has motivated state-of-the-art…

cs.CV2024

Human-in-the-Loop Visual Re-ID for Population Size Estimation

Gustavo Perez, Daniel Sheldon, Grant Van Horn +1

Computer vision-based re-identification (Re-ID) systems are increasingly being deployed for estimating population size in large image collections. However, the estimated size can b…

stat.ME2021

Sibling Regression for Generalized Linear Models

Shiv Shankar, Daniel Sheldon

Field observations form the basis of many scientific studies, especially in ecological and social sciences. Despite efforts to conduct such surveys in a standardized way, observati…

cs.LG2016

Consistently Estimating Markov Chains with Noisy Aggregate Data

Garrett Bernstein, Daniel Sheldon

We address the problem of estimating the parameters of a time-homogeneous Markov chain given only noisy, aggregate data. This arises when a population of individuals behave indepen…

cs.LG2024

Private Regression via Data-Dependent Sufficient Statistic Perturbation

Cecilia Ferrando, Daniel Sheldon

Sufficient statistic perturbation (SSP) is a widely used method for differentially private linear regression. SSP adopts a data-independent approach where privacy noise from a simp…

stat.ML2016

Bethe Projections for Non-Local Inference

Luke Vilnis, David Belanger, Daniel Sheldon +1

Many inference problems in structured prediction are naturally solved by augmenting a tractable dependency structure with complex, non-local auxiliary objectives. This includes the…

cs.CV2021

The Spatio-Temporal Poisson Point Process: A Simple Model for the Alignment of Event Camera Data

Cheng Gu, Erik Learned-Miller, Daniel Sheldon +2

Event cameras, inspired by biological vision systems, provide a natural and data efficient representation of visual information. Visual information is acquired in the form of event…

cs.CR2021

Winning the NIST Contest: A scalable and general approach to differentially private synthetic data

Ryan McKenna, Gerome Miklau, Daniel Sheldon

We propose a general approach for differentially private synthetic data generation, that consists of three steps: (1) select a collection of low-dimensional marginals, (2) measure…

cs.CV2025

Active Measurement: Efficient Estimation at Scale

Max Hamilton, Jinlin Lai, Wenlong Zhao +2

AI has the potential to transform scientific discovery by analyzing vast datasets with little human effort. However, current workflows often do not provide the accuracy or statisti…

cs.LG2021

Faster Kernel Interpolation for Gaussian Processes

Mohit Yadav, Daniel Sheldon, Cameron Musco

A key challenge in scaling Gaussian Process (GP) regression to massive datasets is that exact inference requires computation with a dense n x n kernel matrix, where n is the number…

stat.ML2022

Variational Marginal Particle Filters

Jinlin Lai, Justin Domke, Daniel Sheldon

Variational inference for state space models (SSMs) is known to be hard in general. Recent works focus on deriving variational objectives for SSMs from unbiased sequential Monte Ca…

cs.SI2012

Maximizing the Spread of Cascades Using Network Design

Daniel Sheldon, Bistra Dilkina, Adam N. Elmachtoub +8

We introduce a new optimization framework to maximize the expected spread of cascades in networks. Our model allows a rich set of actions that directly manipulate cascade dynamics…

cs.LG2025

Consensus-Driven Active Model Selection

Justin Kay, Grant Van Horn, Subhransu Maji +2

The widespread availability of off-the-shelf machine learning models poses a challenge: which model, of the many available candidates, should be chosen for a given data analysis ta…

cs.LG2024

Efficient and Private Marginal Reconstruction with Local Non-Negativity

Brett Mullins, Miguel Fuentes, Yingtai Xiao +3

Differential privacy is the dominant standard for formal and quantifiable privacy and has been used in major deployments that impact millions of people. Many differentially private…

cs.AI2016

Robust Optimization for Tree-Structured Stochastic Network Design

Xiaojian Wu, Akshat Kumar, Daniel Sheldon +1

Stochastic network design is a general framework for optimizing network connectivity. It has several applications in computational sustainability including spatial conservation pla…

cs.CR2020

Permute-and-Flip: A new mechanism for differentially private selection

Ryan McKenna, Daniel Sheldon

We consider the problem of differentially private selection. Given a finite set of candidate items and a quality score for each item, our goal is to design a differentially private…

cs.CV2026

Active Measurement of Two-Point Correlations

Max Hamilton, Daniel Sheldon, Subhransu Maji

Two-point correlation functions (2PCF) are widely used to characterize how points cluster in space. In this work, we study the problem of measuring the 2PCF over a large set of poi…

cs.CV2026

Scalable Model-Assisted Multi-Target Estimation in Large Image Collections

Max Hamilton, Jinlin Lai, Daniel Sheldon +1

Computer vision models are increasingly used as measurement tools to estimate population-level quantities from large image collections, but prediction errors introduce bias and the…

cs.LG2019

Differentially Private Bayesian Linear Regression

Garrett Bernstein, Daniel Sheldon

Linear regression is an important tool across many fields that work with sensitive human-sourced data. Significant prior work has focused on producing differentially private point…

cs.LG2023

Sample Average Approximation for Black-Box VI

Javier Burroni, Justin Domke, Daniel Sheldon

We present a novel approach for black-box VI that bypasses the difficulties of stochastic gradient ascent, including the task of selecting step-sizes. Our approach involves using a…

cs.LG2019

Graphical-model based estimation and inference for differential privacy

Ryan McKenna, Daniel Sheldon, Gerome Miklau

Many privacy mechanisms reveal high-level information about a data distribution through noisy measurements. It is common to use this information to estimate the answers to new quer…

cs.LG2023

Automatically Marginalized MCMC in Probabilistic Programming

Jinlin Lai, Javier Burroni, Hui Guan +1

Hamiltonian Monte Carlo (HMC) is a powerful algorithm to sample latent variables from Bayesian models. The advent of probabilistic programming languages (PPLs) frees users from wri…

cs.DB2026

Fast Private Adaptive Query Answering for Large Data Domains

Miguel Fuentes, Brett Mullins, Yingtai Xiao +3

Privately releasing marginals of a tabular dataset is a foundational problem in differential privacy. However, state-of-the-art mechanisms suffer from a computational bottleneck wh…

stat.ML2018

Learning in Integer Latent Variable Models with Nested Automatic Differentiation

Daniel Sheldon, Kevin Winner, Debora Sujono

We develop nested automatic differentiation (AD) algorithms for exact inference and learning in integer latent variable models. Recently, Winner, Sujono, and Sheldon showed how to…

cs.LG2020

Advances in Black-Box VI: Normalizing Flows, Importance Weighting, and Optimization

Abhinav Agrawal, Daniel Sheldon, Justin Domke

Recent research has seen several advances relevant to black-box VI, but the current state of automatic posterior inference is unclear. One such advance is the use of normalizing fl…

cs.SI2013

Collective Diffusion Over Networks: Models and Inference

Akshat Kumar, Daniel Sheldon, Biplav Srivastava

Diffusion processes in networks are increasingly used to model the spread of information and social influence. In several applications in computational sustainability such as the s…

cs.LG2024

Joint Selection: Adaptively Incorporating Public Information for Private Synthetic Data

Miguel Fuentes, Brett Mullins, Ryan McKenna +2

Mechanisms for generating differentially private synthetic data based on marginals and graphical models have been successful in a wide range of settings. However, one limitation of…

cs.LG2021

Parametric Bootstrap for Differentially Private Confidence Intervals

Cecilia Ferrando, Shufan Wang, Daniel Sheldon

The goal of this paper is to develop a practical and general-purpose approach to construct confidence intervals for differentially private parametric estimation. We find that the p…

cs.LG2021

Relaxed Marginal Consistency for Differentially Private Query Answering

Ryan McKenna, Siddhant Pradhan, Daniel Sheldon +1

Many differentially private algorithms for answering database queries involve a step that reconstructs a discrete data distribution from noisy measurements. This provides consisten…

cs.LG2018

Differentially Private Bayesian Inference for Exponential Families

Garrett Bernstein, Daniel Sheldon

The study of private inference has been sparked by growing concern regarding the analysis of data when it stems from sensitive sources. We present the first method for private Baye…