activity
20232026
collaborators

9 papers

stat.ME2026

Bayesian nonparametric inference for modal missing species and features

Alessandro Colombi, Mario Beraha, Daniele Durante +1

Species and feature sampling problems arise naturally whenever each observed unit is associated with one or more labels from a countable alphabet, and inference focuses on the unob…

stat.ME2026

Confidence intervals for maximum unseen probabilities, with application to sequential sampling design

Alessandro Colombi, Mario Beraha, Amichai Painsky +1

Discovery problems often require deciding whether additional sampling is needed to detect all categories whose prevalence exceeds a prespecified threshold. We study this question u…

stat.ML2025

Conformal Inference for Open-Set and Imbalanced Classification

Tianmin Xie, Yanfei Zhou, Ziyi Liang +2

This paper presents a conformal prediction method for classification in highly imbalanced and open-set settings, where there are many possible classes and not all may be represente…

stat.ME2025

Large-scale entity resolution via microclustering Ewens--Pitman random partitions

Mario Beraha, Stefano Favaro

We introduce the microclustering Ewens--Pitman model for random partitions, obtained by scaling the strength parameter of the Ewens--Pitman model linearly with the sample size. The…

stat.AP2025

Online activity prediction via generalized Indian buffet process models

Mario Beraha, Lorenzo Masoero, Stefano Favaro +1

Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism:…

stat.ME2024

Improved prediction of future user activity in online A/B testing

Lorenzo Masoero, Mario Beraha, Thomas Richardson +1

In online randomized experiments or A/B tests, accurate predictions of participant inclusion rates are of paramount importance. These predictions not only guide experimenters in op…