Publications (19)
A Noninformative Prior on a Space of Distribution Functions
Alexander Terenin, David Draper
In a given problem, the Bayesian statistical paradigm requires the specification of a prior distribution that quantifies relevant information about the unknowns of main interest ex…
Discussion of Martingale Posterior Distributions by E. Fong, C. Holmes, and S. G. Walker
David Draper, Erdong Guo
In this discussion note, we respond to the fascinating paper "Martingale Posterior Distributions" by E. Fong, C. Holmes, and S. G. Walker with a couple of comments. On the basis of…
Cox's Theorem and the Jaynesian Interpretation of Probability
Alexander Terenin, David Draper
There are multiple proposed interpretations of probability theory: one such interpretation is true-false logic under uncertainty. Cox's Theorem is a representation theorem that sta…
Pólya Urn Latent Dirichlet Allocation: a doubly sparse massively parallel sampler
Alexander Terenin, MÃ¥ns Magnusson, Leif Jonsson +1
Latent Dirichlet Allocation (LDA) is a topic model widely used in natural language processing and machine learning. Most approaches to training the model rely on iterative algorith…
The Practical Scope of the Central Limit Theorem
David Draper, Erdong Guo
The \textit{Central Limit Theorem (CLT)} is at the heart of a great deal of applied problem-solving in statistics and data science, but the theorem is silent on an important implem…
Infinitely Wide Tensor Networks as Gaussian Process
Erdong Guo, David Draper
Gaussian Process is a non-parametric prior which can be understood as a distribution on the function space intuitively. It is known that by introducing appropriate prior to the wei…
A Space-Based Observational Strategy for Characterizing the First Stars and Galaxies Using the Redshifted 21-cm Global Spectrum
Jack O. Burns, Richard Bradley, Keith Tauscher +19
The redshifted 21-cm monopole is expected to be a powerful probe of the epoch of the first stars and galaxies (). The global 21-cm signal is sensitive to the thermal and i…
A nonparametric Bayesian analysis of heterogeneous treatment effects in digital experimentation
Matt Taddy, Matt Gardner, Liyun Chen +1
Randomized controlled trials play an important role in how Internet companies predict the impact of policy decisions and product changes. In these `digital experiments', different…
Causal Inference in Repeated Observational Studies: A Case Study of eBay Product Releases
Vadim von Brzeski, Matt Taddy, David Draper
Causal inference in observational studies is notoriously difficult, due to the fact that the experimenter is not in charge of the treatment assignment mechanism. Many potential con…
Annealing Double-Head: An Architecture for Online Calibration of Deep Neural Networks
Erdong Guo, David Draper, Maria De Iorio
Model calibration, which is concerned with how frequently the model predicts correctly, not only plays a vital part in statistical model design, but also has substantial practical…
A Simple Necessary Condition For Independence of Real-Valued Random Variables
David Draper, Erdong Guo, Robert Lund +1
The standard method to check for the independence of two real-valued random variables -- demonstrating that the bivariate joint distribution factors into the product of its margina…
GPU-accelerated Gibbs sampling: a case study of the Horseshoe Probit model
Alexander Terenin, Shawfeng Dong, David Draper
Gibbs sampling is a widely used Markov chain Monte Carlo (MCMC) method for numerically approximating integrals of interest in Bayesian statistics and other mathematical sciences. M…
Representation Theorem for Matrix Product States
Erdong Guo, David Draper
In this work, we investigate the universal representation capacity of the Matrix Product States (MPS) from the perspective of boolean functions and continuous functions. We show th…
A Unified Framework for Cluster Methods with Tensor Networks
Erdong Guo, David Draper
Markov Chain Monte Carlo (MCMC), and Tensor Networks (TN) are two powerful frameworks for numerically investigating many-body systems, each offering distinct advantages. MCMC, with…
Comment: A brief survey of the current state of play for Bayesian computation in data science at Big-Data scale
David Draper, Alexander Terenin
We wish to contribute to the discussion of "Comparing Consensus Monte Carlo Strategies for Distributed Bayesian Computation" by offering our views on the current best methods for B…
Asynchronous Gibbs Sampling
Alexander Terenin, Daniel Simpson, David Draper
Gibbs sampling is a Markov Chain Monte Carlo (MCMC) method often used in Bayesian learning. MCMC methods can be difficult to deploy on parallel and distributed systems due to their…
Neural Tangent Kernel of Matrix Product States: Convergence and Applications
Erdong Guo, David Draper
In this work, we study the Neural Tangent Kernel (NTK) of Matrix Product States (MPS) and the convergence of its NTK in the infinite bond dimensional limit. We prove that the NTK o…
Power-Expected-Posterior Priors for Variable Selection in Gaussian Linear Models
Dimitris Fouskakis, Ioannis Ntzoufras, David Draper
In the context of the expected-posterior prior (EPP) approach to Bayesian variable selection in linear models, we combine ideas from power-prior and unit-information-prior methodol…
The Bayesian Method of Tensor Networks
Erdong Guo, David Draper
Bayesian learning is a powerful learning framework which combines the external information of the data (background information) with the internal information (training data) in a l…