4 papers
Sensitivity Sampling with Predictions for k-Means Clustering
Cristian Boldrin, Fabio Vandin
We study the problem of k-means clustering on large datasets. The state-of-the-art for the problem is given by coresets-based approaches, which build small weighted summaries of th…
Scalable and Distributed Silhouette Approximation
Ilie Sarpe, Federico Altieri, Andrea Pietracaprina +2
The silhouette is one of the most widely used measures to assess the quality of a -clustering of a dataset of elements. Its evaluation requires no information beyond the clu…
Few-Shot Resampling for Scalable Statistically-Sound Data Mining
Leonardo Pellegrina, Fabio Vandin
A key step in knowledge discovery is the evaluation of data mining results. In several applications, including pattern mining, graph analysis, and others, this step includes the ev…
Efficient Approximate Temporal Triangle Counting in Streaming with Predictions
Giorgio Venturin, Ilie Sarpe, Fabio Vandin
Triangle counting is a fundamental and widely studied problem on static graphs, and recently on temporal graphs, where edges carry information on the timings of the associated even…