Semi-supervised clustering methods
arXiv:1307.0252 · doi:10.1002/wics.1270
Abstract
Cluster analysis methods seek to partition a data set into homogeneous subgroups. It is useful in a wide variety of applications, including document processing and modern genetics. Conventional clustering methods are unsupervised, meaning that there is no outcome variable nor is anything known about the relationship between the observations in the data set. In many situations, however, information about the clusters is available in addition to the values of the features. For example, the cluster labels of some observations may be known, or certain observations may be known to belong to the same cluster. In other cases, one may wish to identify clusters that are associated with a particular outcome variable. This review describes several clustering algorithms (known as "semi-supervised clustering" methods) that can be applied in these situations. The majority of these methods are modifications of the popular k-means clustering method, and several of them will be described in detail. A brief description of some other semi-supervised clustering algorithms is also provided.
28 pages, 5 figures
Cited by in corpus (16)
- Structured Sparse Subspace Clustering: A Joint Affinity Learning and Subspace Clustering Framework
- The Challenge of Non-Technical Loss Detection using Artificial Intelligence: A Survey
- AutoEmbedder: A semi-supervised DNN embedding system for clustering
- Integrating Prior Knowledge in Mixed Initiative Social Network Clustering
- Semi-supervised Clustering for Short Text via Deep Representation Learning
- Semi-supervised Text Categorization Using Recursive K-means Clustering
- Outcome-Driven Clustering of Acute Coronary Syndrome Patients using Multi-Task Neural Network with Attention
- Onset of a conceptual outline map to get a hold on the jungle of cluster analysis
- Hierarchical Clustering with Prior Knowledge
- ConiVAT: Cluster Tendency Assessment and Clustering with Partial Background Knowledge
- Matrix Completion with Prior Subspace Information via Maximizing Correlation
- Constrained Hierarchical Clustering via Graph Coarsening and Optimal Cuts
- Navigating Uncertainties in Machine Learning for Structural Dynamics: A Comprehensive Survey of Probabilistic and Non-Probabilistic Approaches in Forward and Inverse Problems
- Deep Goal-Oriented Clustering
- Outcome-Guided Disease Subtyping for High-Dimensional Omics Data
- Auditing for Diversity using Representative Examples