Simultaneous identification of specifically interacting paralogs and inter-protein contacts by Direct-Coupling Analysis
arXiv:1605.03745 · doi:10.1073/pnas.1607570113
Abstract
Understanding protein-protein interactions is central to our understanding of almost all complex biological processes. Computational tools exploiting rapidly growing genomic databases to characterize protein-protein interactions are urgently needed. Such methods should connect multiple scales from evolutionary conserved interactions between families of homologous proteins, over the identification of specifically interacting proteins in the case of multiple paralogs inside a species, down to the prediction of residues being in physical contact across interaction interfaces. Statistical inference methods detecting residue-residue coevolution have recently triggered considerable progress in using sequence data for quaternary protein structure prediction; they require, however, large joint alignments of homologous protein pairs known to interact. The generation of such alignments is a complex computational task on its own; application of coevolutionary modeling has in turn been restricted to proteins without paralogs, or to bacterial systems with the corresponding coding genes being co-localized in operons. Here we show that the Direct-Coupling Analysis of residue coevolution can be extended to connect the different scales, and simultaneously to match interacting paralogs, to identify inter-protein residue-residue contacts and to discriminate interacting from noninteracting families in a multiprotein system. Our results extend the potential applications of coevolutionary analysis far beyond cases treatable so far.
Main Text 19 pages Supp. Inf. 16 pages
References in corpus (4)
- Identification of direct residue contacts in protein-protein interaction by message passing
- Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
- Inferring interaction partners from protein sequences
- Fast and accurate multivariate Gaussian modeling of protein families: Predicting residue contacts and protein-interaction partners
Cited by in corpus (17)
- Inverse Statistical Physics of Protein Sequences: A Key Issues Review
- Large-scale identification of coevolution signals across homo-oligomeric protein interfaces by Direct Coupling Analysis
- Inter-residue, inter-protein and inter-family coevolution: bridging the scales
- Inferring interaction partners from protein sequences using mutual information
- Generative power of a protein language model trained on multiple sequence alignments
- Phylogenetic correlations can suffice to infer protein partners from sequences
- Revealing evolutionary constraints on proteins through sequence analysis
- Correlations from structure and phylogeny combine constructively in the inference of protein partners from sequences
- Aligning biological sequences by exploiting residue conservation and coevolution
- Correlation-Compressed Direct Coupling Analysis
- Combining phylogeny and coevolution improves the inference of interaction partners among paralogous proteins
- Statistical physics of interacting proteins: impact of dataset size and quality assessed in synthetic sequences
- Inferring genetic fitness from genomic data
- Impact of phylogeny on structural contact inference from protein sequence data
- Improved Pseudolikelihood Regularization and Decimation methods on Non-linearly Interacting Systems with Continuous Variables
- DiffPaSS -- High-performance differentiable pairing of protein sequences using soft scores
- Impact of phylogeny on the inference of functional sectors from protein sequence data