Publications (11)
Doubly Non-Central Beta Matrix Factorization for DNA Methylation Data
Aaron Schein, Anjali Nagulpally, Hanna Wallach +1
We present a new non-negative matrix factorization model for bounded-support data based on the doubly non-central beta (DNCB) distribution, a generalization of the beta dis…
A global optimization algorithm for sparse mixed membership matrix factorization
Fan Zhang, Chuangqi Wang, Andrew Trapp +1
Mixed membership factorization is a popular approach for analyzing data sets that have within-sample heterogeneity. In recent years, several algorithms have been developed for mixe…
Variational inference for rare variant detection in deep, heterogeneous next-generation sequencing data
Fan Zhang, Patrick Flaherty
The detection of rare variants is important for understanding the genetic heterogeneity in mixed samples. Recently, next-generation sequencing (NGS) technologies have enabled the i…
Hedging in Sequential Experiments
Thomas Cook, Patrick Flaherty
Experimentation involves risk. The investigator expends time and money in the pursuit of data that supports a hypothesis. In the end, the investigator may find that all of these co…
Doubly Non-Central Beta Matrix Factorization for Stable Dimensionality Reduction of Bounded Support Matrix Data
Anjali N. Albert, Patrick Flaherty, Aaron Schein
We consider the problem of developing interpretable and computationally efficient matrix decomposition methods for matrices whose entries have bounded support. Such matrices are fo…
Exact and Approximate Hierarchical Clustering Using A*
Craig S. Greenberg, Sebastian Macaluso, Nicholas Monath +6
Hierarchical clustering is a critical task in numerous domains. Many approaches are based on heuristics and the properties of the resulting clusterings are studied post hoc. Howeve…
Maximum a Posteriori Inference for Factor Graphs via Benders' Decomposition
Harsh Vardhan Dubey, Ji Ah Lee, Patrick Flaherty
Many Bayesian statistical inference problems come down to computing a maximum a-posteriori (MAP) assignment of latent variables. Yet, standard methods for estimating the MAP assign…
Cost-aware Generalized -investing for Multiple Hypothesis Testing
Thomas Cook, Harsh Vardhan Dubey, Ji Ah Lee +3
We consider the problem of sequential multiple hypothesis testing with nontrivial data collection costs. This problem appears, for example, when conducting biological experiments t…
Data Structures & Algorithms for Exact Inference in Hierarchical Clustering
Craig S. Greenberg, Sebastian Macaluso, Nicholas Monath +5
Hierarchical clustering is a fundamental task often used to discover meaningful structures in data, such as phylogenetic trees, taxonomies of concepts, subtypes of cancer, and casc…
A Deterministic Global Optimization Method for Variational Inference
Hachem Saddiki, Andrew C. Trapp, Patrick Flaherty
Variational inference methods for latent variable statistical models have gained popularity because they are relatively fast, can handle large data sets, and have deterministic con…
MAP Clustering under the Gaussian Mixture Model via Mixed Integer Nonlinear Optimization
Patrick Flaherty, Pitchaya Wiratchotisatian, Ji Ah Lee +2
We present a global optimization approach for solving the maximum a-posteriori (MAP) clustering problem under the Gaussian mixture model.Our approach can accommodate side constrain…