Identification of direct residue contacts in protein-protein interaction by message passing
arXiv:0901.1248 · doi:10.1073/pnas.0805923106
Abstract
Understanding the molecular determinants of specificity in protein-protein interaction is an outstanding challenge of postgenome biology. The availability of large protein databases generated from sequences of hundreds of bacterial genomes enables various statistical approaches to this problem. In this context covariance-based methods have been used to identify correlation between amino acid positions in interacting proteins. However, these methods have an important shortcoming, in that they cannot distinguish between directly and indirectly correlated residues. We developed a method that combines covariance analysis with global inference analysis, adopted from use in statistical physics. Applied to a set of >2,500 representatives of the bacterial two-component signal transduction system, the combination of covariance with global inference successfully and robustly identified residue pairs that are proximal in space without resorting to ad hoc tuning parameters, both for heterointeractions between sensor kinase (SK) and response regulator (RR) proteins and for homointeractions between RR proteins. The spectacular success of this approach illustrates the effectiveness of the global inference approach in identifying direct interaction based on sequence information alone. We expect this method to be applicable soon to interaction surfaces between proteins present in only 1 copy per genome as the number of sequenced genomes continues to expand. Use of this method could significantly increase the potential targets for therapeutic intervention, shed light on the mechanism of protein-protein interaction, and establish the foundation for the accurate prediction of interacting protein partners.
Supplementary information available on http://www.pnas.org/content/106/1/67.abstract
References in corpus (1)
Cited by in corpus (61)
- Accurate De Novo Prediction of Protein Contact Map by Ultra-Deep Learning Model
- Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
- Fast and accurate multivariate Gaussian modeling of protein families: Predicting residue contacts and protein-interaction partners
- Mean Field Theory For Non-Equilibrium Network Reconstruction
- Protein sectors: statistical coupling analysis versus conservation
- Epistatic models predict mutable sites in SARS-CoV-2 proteins and epitopes
- Improving contact prediction along three dimensions
- Inter-residue, inter-protein and inter-family coevolution: bridging the scales
- Dynamical criticality in the collective activity of a population of retinal neurons
- Protein language models trained on multiple sequence alignments learn phylogenetic relationships
- Influence of Multiple Sequence Alignment Depth on Potts Statistical Models of Protein Covariation
- Benchmarking inverse statistical approaches for protein structure and design with exactly solvable models
- Generative power of a protein language model trained on multiple sequence alignments
- A perspective on protein structure prediction using quantum computers
- Mean-field theory for the inverse Ising problem at low temperatures
- Dynamical TAP equations for non-equilibrium Ising spin glasses
- Computational protein design with evolutionary-based and physics-inspired modeling: current and future synergies
- Short-range interaction vs long-range correlation in bird flocks
- Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
- Large Pseudo-Counts and -Norm Penalties Are Necessary for the Mean-Field Inference of Ising and Potts Models
- Inference of kinetic Ising model on sparse graphs
- Unsupervised hierarchical clustering using the learning dynamics of RBMs
- Memory-free dynamics for the TAP equations of Ising models with arbitrary rotation invariant ensembles of random coupling matrices
- Inverse Ising problem in continuous time: A latent variable approach
- Improving landscape inference by integrating heterogeneous data in the inverse Ising problem
- Correlations from structure and phylogeny combine constructively in the inference of protein partners from sequences
- Inference of Co-Evolving Site Pairs: an Excellent Predictor of Contact Residue Pairs in Protein 3D structures
- Dynamics and Performance of Susceptibility Propagation on Synthetic Data
- From evolution to folding of repeat proteins
- Combining phylogeny and coevolution improves the inference of interaction partners among paralogous proteins
- Unsupervisedly Prompting AlphaFold2 for Few-Shot Learning of Accurate Folding Landscape and Protein Structure Prediction
- A statistical physics approach to learning curves for the Inverse Ising problem
- Network reconstruction via the minimum description length principle
- Inverse Ising inference with correlated samples
- Impact of phylogeny on structural contact inference from protein sequence data
- Genome-Wide Epigenetic Modifications as a Shared Memory Consensus Problem
- Unified framework for modeling multivariate distributions in biological sequences
- Impact of population size on early adaptation in rugged fitness landscapes
- Interpretable machine learning of amino acid patterns in proteins: a statistical ensemble approach
- Network inference in the non-equilibrium steady state
- Biases in Inverse Ising Estimates of Near-Critical Behaviour
- Evolutionary Dynamics of a Lattice Dimer: a Toy Model for Stability vs. Affinity Trade-offs in Proteins
- DiffPaSS -- High-performance differentiable pairing of protein sequences using soft scores
- Phylogenetic Corrections and Higher-Order Sequence Statistics in Protein Families: The Potts Model vs MSA Transformer
- Discovering sparse control strategies in C. elegans
- Inferring Higher-Order Couplings with Neural Networks
- Inverse problems in spin models
- Impact of phylogeny on the inference of functional sectors from protein sequence data
- The cavity method to protein design problem
- Inverse modeling of time-delayed interactions via the dynamic-entropy formalism
- AmoebaContact and GDFold: a new pipeline for rapid prediction of protein structures
- Information costs in the control of protein synthesis
- Unsupervised and Supervised Structure Learning for Protein Contact Prediction
- adabmDCA: Adaptive Boltzmann machine learning for biological sequences
- JUWELS Booster -- A Supercomputer for Large-Scale AI Research
- Machine-Learned Molecular Surface and Its Application to Implicit Solvent Simulation
- Cavity approach for modeling and fitting polymer stretching
- Perturbational Decomposition Analysis for Quantum Ising Model with Weak Transverse Fields
- Classification of SARS-CoV-2 Variants through The Epistatical Circos Plots with Convolutional Neural Networks
- Mapping Inter-City Trade Networks to Maximum Entropy Models using Electronic Invoice Data
- Direct Information Reweighted by Contact Templates: Improved RNA Contact Prediction by Combining Structural Features