Consensus Tree Estimation with False Discovery Control via Partially Ordered Sets
arXiv:2511.23433
Abstract
Trees are data objects that hierarchically organize categories. Collections of trees arise in a diverse variety of fields, including evolutionary biology, machine learning, social sciences and anatomy. Summarizing a collection of trees by a single representative is challenging, in part due to the dimensions of both the sample space and the parameter space. We frame consensus tree estimation as a structured feature-selection problem, where leaves and edges are the features. We introduce a partial order on trees, use it to define false discoveries for a candidate summary tree, and develop novel estimation algorithms that control the false discovery rate at a nominal level for a broad class of generative models. We also use the partial order to assess the stability of features in a selected tree. Importantly, our method accommodates unequal leaf sets and non-binary trees, which commonly arise in modern datasets. Our feature-selection perspective yields finite-sample and model-free guarantees and provides a foundation for integrating multiple testing tools into tree estimation. We apply the method to study the origins of complex life. Our estimated tree recapitulates well-known divisions but highlights that there is insufficient data to determine the most recent archaeal ancestor of eukaryotic life.