From principal component to direct coupling analysis of coevolution in proteins: Low-eigenvalue modes are needed for structure prediction
arXiv:1212.3281 · doi:10.1371/journal.pcbi.1003176
Abstract
Various approaches have explored the covariation of residues in multiple-sequence alignments of homologous proteins to extract functional and structural information. Among those are principal component analysis (PCA), which identifies the most correlated groups of residues, and direct coupling analysis (DCA), a global inference method based on the maximum entropy principle, which aims at predicting residue-residue contacts. In this paper, inspired by the statistical physics of disordered systems, we introduce the Hopfield-Potts model to naturally interpolate between these two approaches. The Hopfield-Potts model allows us to identify relevant 'patterns' of residues from the knowledge of the eigenmodes and eigenvalues of the residue-residue correlation matrix. We show how the computation of such statistical patterns makes it possible to accurately predict residue-residue contacts with a much smaller number of parameters than DCA. This dimensional reduction allows us to avoid overfitting and to extract contact information from multiple-sequence alignments of reduced size. In addition, we show that low-eigenvalue correlation modes, discarded by PCA, are important to recover structural information: the corresponding patterns are highly localized, that is, they are concentrated in few sites, which we find to be in close contact in the three-dimensional protein fold.
Supporting information can be downloaded from: http://www.ploscompbiol.org/article/info:doi/10.1371/journal.pcbi.1003176
References in corpus (3)
Cited by in corpus (31)
- Fast pseudolikelihood maximization for direct-coupling analysis of protein structure from many homologous amino-acid sequences
- Fast and accurate multivariate Gaussian modeling of protein families: Predicting residue contacts and protein-interaction partners
- Epigenetic landscapes explain partially reprogrammed cells and identify key reprogramming genes
- Improving contact prediction along three dimensions
- On the sufficiency of pairwise interactions in maximum entropy models of biological networks
- Orthogonal Sparse PCA and Covariance Estimation via Procrustes Reformulation
- Selection of sequence motifs and generative Hopfield-Potts models for protein familiesilies
- Phylogenetic correlations can suffice to infer protein partners from sequences
- Beyond position weight matrices: nucleotide correlations in transcription factor binding sites and their description
- Revealing evolutionary constraints on proteins through sequence analysis
- Parsimonious evolutionary scenario for the origin of allostery and coevolution patterns in proteins
- Large Pseudo-Counts and -Norm Penalties Are Necessary for the Mean-Field Inference of Ising and Potts Models
- On the entropy of protein families
- Learning Maximum Entropy Models from finite size datasets: a fast Data-Driven algorithm allows sampling from the posterior distribution
- Statistical Physics and Representations in Real and Artificial Neural Networks
- Identifying relevant positions in proteins by Critical Variable Selection
- Estimating the principal components of correlation matrices from all their empirical eigenvectors
- Bézier interpolation improves the inference of dynamical models from data
- A statistical-inference approach to reconstruct inter-cellular interactions in cell-migration experiments
- Asymptotics of eigenstructure of sample correlation matrices for high-dimensional spiked models
- Inverse Ising inference with correlated samples
- Efficient training of energy-based models via spin-glass control
- Progress and open problems in evolutionary dynamics
- Descriptive power and predictive limits of a discrete Hasimoto--DNLS model of protein backbone structure
- Worldwide bilateral geopolitical interactions network inferred from national disciplinary profiles
- Impact of phylogeny on the inference of functional sectors from protein sequence data
- Unsupervised and Supervised Structure Learning for Protein Contact Prediction
- Multidimensional mutual information methods for the analysis of covariation in multiple sequence alignments
- A FEAST variant incorporated with a power iteration
- Mutational dynamics of influenza A viruses: a principal component analysis of hemagglutinin sequences of subtype H1
- Inference of principal components of noisy correlation matrices with prior information