Inferring interaction partners from protein sequences
arXiv:1604.08354 · doi:10.1073/pnas.1606762113
Abstract
Specific protein-protein interactions are crucial in the cell, both to ensure the formation and stability of multi-protein complexes, and to enable signal transduction in various pathways. Functional interactions between proteins result in coevolution between the interaction partners, causing their sequences to be correlated. Here we exploit these correlations to accurately identify which proteins are specific interaction partners from sequence data alone. Our general approach, which employs a pairwise maximum entropy model to infer couplings between residues, has been successfully used to predict the three-dimensional structures of proteins from sequences. Thus inspired, we introduce an iterative algorithm to predict specific interaction partners from two protein families whose members are known to interact. We first assess the algorithm's performance on histidine kinases and response regulators from bacterial two-component signaling systems. We obtain a striking 0.93 true positive fraction on our complete dataset without any a priori knowledge of interaction partners, and we uncover the origin of this success. We then apply the algorithm to proteins from ATP-binding cassette (ABC) transporter complexes, and obtain accurate predictions in these systems as well. Finally, we present two metrics that accurately distinguish interacting protein families from non-interacting ones, using only sequence data.
25 pages, 19 figures, published version
References in corpus (4)
- Identification of direct residue contacts in protein-protein interaction by message passing
- Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
- Fast and accurate multivariate Gaussian modeling of protein families: Predicting residue contacts and protein-interaction partners
- Benchmarking inverse statistical approaches for protein structure and design with exactly solvable models
Cited by in corpus (21)
- Inverse Statistical Physics of Protein Sequences: A Key Issues Review
- Simultaneous identification of specifically interacting paralogs and inter-protein contacts by Direct-Coupling Analysis
- Large-scale identification of coevolution signals across homo-oligomeric protein interfaces by Direct Coupling Analysis
- Inter-residue, inter-protein and inter-family coevolution: bridging the scales
- Inferring interaction partners from protein sequences using mutual information
- Generative power of a protein language model trained on multiple sequence alignments
- Phylogenetic correlations can suffice to infer protein partners from sequences
- Revealing evolutionary constraints on proteins through sequence analysis
- Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
- Quantifying Relevance in Learning and Inference
- Aligning biological sequences by exploiting residue conservation and coevolution
- Correlations from structure and phylogeny combine constructively in the inference of protein partners from sequences
- Building blocks of protein structures -- Physics meets Biology
- Combining phylogeny and coevolution improves the inference of interaction partners among paralogous proteins
- Statistical physics of interacting proteins: impact of dataset size and quality assessed in synthetic sequences
- On generative models of T-cell receptor sequences
- Impact of phylogeny on structural contact inference from protein sequence data
- DiffPaSS -- High-performance differentiable pairing of protein sequences using soft scores
- Impact of phylogeny on the inference of functional sectors from protein sequence data
- Uncovering the non-equilibrium stationary properties in sparse Boolean networks
- Unsupervised Bayesian Ising Approximation for revealing the neural dictionary in songbirds