papers

Publications (11)

stat.ML2021

Doubly Non-Central Beta Matrix Factorization for DNA Methylation Data

Aaron Schein, Anjali Nagulpally, Hanna Wallach +1

We present a new non-negative matrix factorization model for bounded-support data based on the doubly non-central beta (DNCB) distribution, a generalization of the beta dis…

stat.ME2016

A global optimization algorithm for sparse mixed membership matrix factorization

Fan Zhang, Chuangqi Wang, Andrew Trapp +1

Mixed membership factorization is a popular approach for analyzing data sets that have within-sample heterogeneity. In recent years, several algorithms have been developed for mixe…

q-bio.GN2016

Variational inference for rare variant detection in deep, heterogeneous next-generation sequencing data

Fan Zhang, Patrick Flaherty

The detection of rare variants is important for understanding the genetic heterogeneity in mixed samples. Recently, next-generation sequencing (NGS) technologies have enabled the i…

q-fin.RM2024

Hedging in Sequential Experiments

Thomas Cook, Patrick Flaherty

Experimentation involves risk. The investigator expends time and money in the pursuit of data that supports a hypothesis. In the end, the investigator may find that all of these co…

cs.LG2024

Doubly Non-Central Beta Matrix Factorization for Stable Dimensionality Reduction of Bounded Support Matrix Data

Anjali N. Albert, Patrick Flaherty, Aaron Schein

We consider the problem of developing interpretable and computationally efficient matrix decomposition methods for matrices whose entries have bounded support. Such matrices are fo…

cs.LG2021

Exact and Approximate Hierarchical Clustering Using A*

Craig S. Greenberg, Sebastian Macaluso, Nicholas Monath +6

Hierarchical clustering is a critical task in numerous domains. Many approaches are based on heuristics and the properties of the resulting clusterings are studied post hoc. Howeve…

stat.ML2024

Maximum a Posteriori Inference for Factor Graphs via Benders' Decomposition

Harsh Vardhan Dubey, Ji Ah Lee, Patrick Flaherty

Many Bayesian statistical inference problems come down to computing a maximum a-posteriori (MAP) assignment of latent variables. Yet, standard methods for estimating the MAP assign…

cs.LG2023

Cost-aware Generalized -investing for Multiple Hypothesis Testing

Thomas Cook, Harsh Vardhan Dubey, Ji Ah Lee +3

We consider the problem of sequential multiple hypothesis testing with nontrivial data collection costs. This problem appears, for example, when conducting biological experiments t…

cs.DS2020

Data Structures & Algorithms for Exact Inference in Hierarchical Clustering

Craig S. Greenberg, Sebastian Macaluso, Nicholas Monath +5

Hierarchical clustering is a fundamental task often used to discover meaningful structures in data, such as phylogenetic trees, taxonomies of concepts, subtypes of cancer, and casc…

stat.ME2017

A Deterministic Global Optimization Method for Variational Inference

Hachem Saddiki, Andrew C. Trapp, Patrick Flaherty

Variational inference methods for latent variable statistical models have gained popularity because they are relatively fast, can handle large data sets, and have deterministic con…

stat.ML2020

MAP Clustering under the Gaussian Mixture Model via Mixed Integer Nonlinear Optimization

Patrick Flaherty, Pitchaya Wiratchotisatian, Ji Ah Lee +2

We present a global optimization approach for solving the maximum a-posteriori (MAP) clustering problem under the Gaussian mixture model.Our approach can accommodate side constrain…