collaborators

5 papers

stat.AP2026

Online activity prediction via generalized Indian buffet process models

Mario Beraha, Lorenzo Masoero, Stefano Favaro +1

Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism:…

stat.ME2026

Confidence intervals for maximum unseen probabilities, with application to sequential sampling design

Alessandro Colombi, Mario Beraha, Amichai Painsky +1

Discovery problems often require deciding whether additional sampling is needed to detect all categories whose prevalence exceeds a prespecified threshold. We study this question u…

stat.ML2025

Conformal Inference for Open-Set and Imbalanced Classification

Tianmin Xie, Yanfei Zhou, Ziyi Liang +2

This paper presents a conformal prediction method for classification in highly imbalanced and open-set settings, where there are many possible classes and not all may be represente…

stat.ME2025

Large-scale entity resolution via microclustering Ewens--Pitman random partitions

Mario Beraha, Stefano Favaro

We introduce the microclustering Ewens--Pitman model for random partitions, obtained by scaling the strength parameter of the Ewens--Pitman model linearly with the sample size. The…

stat.ME2025

A smoothed-Bayesian approach to frequency recovery from sketched data

Mario Beraha, Stefano Favaro, Matteo Sesia

We provide a novel statistical perspective on a classical problem at the intersection of computer science and information theory: recovering the empirical frequency of a symbol in…