How pairwise coevolutionary models capture the collective residue variability in proteins
arXiv:1801.04184 · doi:10.1093/molbev/msy007
Abstract
Global coevolutionary models of homologous protein families, as constructed by direct coupling analysis (DCA), have recently gained popularity in particular due to their capacity to accurately predict residue-residue contacts from sequence information alone, and thereby to facilitate tertiary and quaternary protein structure prediction. More recently, they have also been used to predict fitness effects of amino-acid substitutions in proteins, and to predict evolutionary conserved protein-protein interactions. These models are based on two currently unjustified hypotheses: (a) correlations in the amino-acid usage of different positions are resulting collectively from networks of direct couplings; and (b) pairwise couplings are sufficient to capture the amino-acid variability. Here we propose a highly precise inference scheme based on Boltzmann-machine learning, which allows us to systematically address these hypotheses. We show how correlations are built up in a highly collective way by a large number of coupling paths, which are based on the protein's three-dimensional structure. We further find that pairwise coevolutionary models capture the collective residue variability across homologous proteins even for quantities which are not imposed by the inference procedure, like three-residue correlations, the clustered structure of protein families in sequence space or the sequence distances between homologs. These findings strongly suggest that pairwise coevolutionary models are actually sufficient to accurately capture the residue variability in homologous protein families.
17 pages, 3 figures and one table
References in corpus (5)
- Identification of direct residue contacts in protein-protein interaction by message passing
- Accurate De Novo Prediction of Protein Contact Map by Ultra-Deep Learning Model
- Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
- Inferring fitness landscapes by regression produces biased estimates of epistasis
- Inter-residue, inter-protein and inter-family coevolution: bridging the scales
Cited by in corpus (24)
- Efficient generative modeling of protein sequences using simple autoregressive models
- Protein language models trained on multiple sequence alignments learn phylogenetic relationships
- Generative power of a protein language model trained on multiple sequence alignments
- Selection of sequence motifs and generative Hopfield-Potts models for protein familiesilies
- Generative Capacity of Probabilistic Protein Sequence Models
- Parsimonious evolutionary scenario for the origin of allostery and coevolution patterns in proteins
- Sparse generative modeling via parameter-reduction of Boltzmann machines: application to protein-sequence families
- Modeling sequence-space exploration and emergence of epistatic signals in protein evolution
- Emergent time scales of epistasis in protein evolution
- Assessing the accuracy of direct-coupling analysis for RNA contact prediction
- Direct Coupling Analysis of Epistasis in Allosteric Materials
- Correlations from structure and phylogeny combine constructively in the inference of protein partners from sequences
- Is novelty predictable?
- Aligning biological sequences by exploiting residue conservation and coevolution
- Towards Parsimonious Generative Modeling of RNA Families
- Pre-training Co-evolutionary Protein Representation via A Pairwise Masked Language Model
- Impact of phylogeny on structural contact inference from protein sequence data
- Inference of stochastic time series with missing data
- Fluctuations and the limit of predictability in protein evolution
- Boltzmann machine learning and regularization methods for inferring evolutionary fields and couplings from a multiple sequence alignment
- adabmDCA 2.0 -- a flexible but easy-to-use package for Direct Coupling Analysis
- Pseudo-likelihood produces associative memories able to generalize, even for asymmetric couplings
- DCAlign v1.0: Aligning biological sequences using co-evolution models and informed priors
- adabmDCA: Adaptive Boltzmann machine learning for biological sequences