Coupled Two-Way Clustering Analysis of Gene Microarray Data
arXiv:physics/0004009 · doi:10.1073/pnas.210134797
Abstract
We present a novel coupled two-way clustering approach to gene microarray data analysis. The main idea is to identify subsets of the genes and samples, such that when one of these is used to cluster the other, stable and significant partitions emerge. The search for such subsets is a computationally complex task: we present an algorithm, based on iterative clustering, which performs such a search. This analysis is especially suitable for gene microarray data, where the contributions of a variety of biological mechanisms to the gene expression levels are entangled in a large body of experimental data. The method was applied to two gene microarray data sets, on colon cancer and leukemia. By identifying relevant subsets of the data and focusing on them we were able to discover partitions and correlations that were masked and hidden when the full dataset was used in the analysis. Some of these partitions have clear biological interpretation; others can serve to identify possible directions for future research.
References in corpus (1)
Cited by in corpus (30)
- The Iterative Signature Algorithm for the analysis of large scale gene expression data
- Dynamic modeling of gene expression data
- Finding large average submatrices in high dimensional data
- Expression profiles of acute lymphoblastic and myeloblastic leukemias with ALL-1 rearrangements
- Reinforcement Learning based Recommender System using Biclustering Technique
- Algorithms of maximum likelihood data clustering with applications
- Robust Detection of Hierarchical Communities from Escherichia coli Gene Expression Data
- A statistical framework for the analysis of microarray probe-level data
- Helices 2 and 3 are the initiation sites in the PrPc -> PrPsc transition
- Human cancers over express genes that are specific to a variety of normal human tissues
- Link between allosteric signal transduction and functional dynamics in a multi-subunit enzyme: S-adenosylhomocysteine hydrolase
- SUBIC: A Supervised Bi-Clustering Approach for Precision Medicine
- Genomics as a Service: a Joint Computing and Networking Perspective
- Evolutionary Biclustering of Clickstream Data
- Induction in myeloid leukemic cells of genes that are expressed in different normal tissues
- Exact Clustering in Tensor Block Model: Statistical Optimality and Computational Limit
- Profile Likelihood Biclustering
- Probabilistic analysis of the human transcriptome with side information
- Mapping Energy Landscapes of Non-Convex Learning Problems
- A method for visual identification of small sample subgroups and potential biomarkers
- Kernel method for clustering based on optimal target vector
- Algorithm for Finding Optimal Gene Sets in Microarray Prediction
- Biclustering with Alternating K-Means
- Microarray Data Management. An Enterprise Information Approach: Implementations and Challenges
- BAREB: A Bayesian repulsive biclustering model for periodontal data
- Co-clustering of Fuzzy Lagged Data
- Allosteric communication in Dihydrofolate Reductase: Signaling network and pathways for closed to occluded transition and back
- Coupled Two-Way Clustering Analysis of Breast Cancer and Colon Cancer Gene Expression Data
- Biclustering Via Sparse Clustering
- Block clustering with collapsed latent block models