9 papers
Bayesian nonparametric inference for modal missing species and features
Alessandro Colombi, Mario Beraha, Daniele Durante +1
Species and feature sampling problems arise naturally whenever each observed unit is associated with one or more labels from a countable alphabet, and inference focuses on the unob…
Confidence intervals for maximum unseen probabilities, with application to sequential sampling design
Alessandro Colombi, Mario Beraha, Amichai Painsky +1
Discovery problems often require deciding whether additional sampling is needed to detect all categories whose prevalence exceeds a prespecified threshold. We study this question u…
Conformal Inference for Open-Set and Imbalanced Classification
Tianmin Xie, Yanfei Zhou, Ziyi Liang +2
This paper presents a conformal prediction method for classification in highly imbalanced and open-set settings, where there are many possible classes and not all may be represente…
Large-scale entity resolution via microclustering Ewens--Pitman random partitions
Mario Beraha, Stefano Favaro
We introduce the microclustering Ewens--Pitman model for random partitions, obtained by scaling the strength parameter of the Ewens--Pitman model linearly with the sample size. The…
Online activity prediction via generalized Indian buffet process models
Mario Beraha, Lorenzo Masoero, Stefano Favaro +1
Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism:…
Improved prediction of future user activity in online A/B testing
Lorenzo Masoero, Mario Beraha, Thomas Richardson +1
In online randomized experiments or A/B tests, accurate predictions of participant inclusion rates are of paramount importance. These predictions not only guide experimenters in op…