Coevolutionary landscape inference and the context-dependence of mutations in beta-lactamase TEM-1
arXiv:1510.03224 · doi:10.1093/molbev/msv211
Abstract
The quantitative characterization of mutational landscapes is a task of outstanding importance in evolutionary and medical biology: It is, e.g., of central importance for our understanding of the phenotypic effect of mutations related to disease and antibiotic drug resistance. Here we develop a novel inference scheme for mutational landscapes, which is based on the statistical analysis of large alignments of homologs of the protein of interest. Our method is able to capture epistatic couplings between residues, and therefore to assess the dependence of mutational effects on the sequence context where they appear. Compared to recent large-scale mutagenesis data of the beta-lactamase TEM-1, a protein providing resistance against beta-lactam antibiotics, our method leads to an increase of about 40% in explicative power as compared to approaches neglecting epistasis. We find that the informative sequence context extends to residues at native distances of about 20Å from the mutated site, reaching thus far beyond residues in direct physical contact.
14 pages, 5 figures. Supplementary files on the publisher's website: http://mbe.oxfordjournals.org/content/early/2015/10/06/molbev.msv211.short?rss=1
References in corpus (5)
- Identification of direct residue contacts in protein-protein interaction by message passing
- Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
- Fast and accurate multivariate Gaussian modeling of protein families: Predicting residue contacts and protein-interaction partners
- Inferring fitness landscapes by regression produces biased estimates of epistasis
- Large Pseudo-Counts and -Norm Penalties Are Necessary for the Mean-Field Inference of Ising and Potts Models
Cited by in corpus (44)
- Quantification of the effect of mutations using a global probability model of natural sequence variation
- Inverse Statistical Physics of Protein Sequences: A Key Issues Review
- Inferring interaction partners from protein sequences
- Large-scale identification of coevolution signals across homo-oligomeric protein interfaces by Direct Coupling Analysis
- Efficient generative modeling of protein sequences using simple autoregressive models
- Epistatic models predict mutable sites in SARS-CoV-2 proteins and epitopes
- Inferring interaction partners from protein sequences using mutual information
- Influence of Multiple Sequence Alignment Depth on Potts Statistical Models of Protein Covariation
- Benchmarking inverse statistical approaches for protein structure and design with exactly solvable models
- Generative power of a protein language model trained on multiple sequence alignments
- Selection of sequence motifs and generative Hopfield-Potts models for protein familiesilies
- Generative Capacity of Probabilistic Protein Sequence Models
- Phylogenetic correlations can suffice to infer protein partners from sequences
- Computational protein design with evolutionary-based and physics-inspired modeling: current and future synergies
- Revealing evolutionary constraints on proteins through sequence analysis
- Parsimonious evolutionary scenario for the origin of allostery and coevolution patterns in proteins
- Modeling sequence-space exploration and emergence of epistatic signals in protein evolution
- Sparse generative modeling via parameter-reduction of Boltzmann machines: application to protein-sequence families
- Emergent time scales of epistasis in protein evolution
- Inference of compressed Potts graphical models
- Size and structure of the sequence space of repeat proteins
- Direct Coupling Analysis of Epistasis in Allosteric Materials
- Population-specific design of de-immunized protein biotherapeutics
- Random versus maximum entropy models of neural population activity
- Improving landscape inference by integrating heterogeneous data in the inverse Ising problem
- Aligning biological sequences by exploiting residue conservation and coevolution
- Correlations from structure and phylogeny combine constructively in the inference of protein partners from sequences
- Is novelty predictable?
- From evolution to folding of repeat proteins
- Correlation-Compressed Direct Coupling Analysis
- Statistical physics of interacting proteins: impact of dataset size and quality assessed in synthetic sequences
- Inferring genetic fitness from genomic data
- Towards Parsimonious Generative Modeling of RNA Families
- Pre-training Co-evolutionary Protein Representation via A Pairwise Masked Language Model
- Impact of phylogeny on structural contact inference from protein sequence data
- Statistical Genetics in and out of Quasi-Linkage Equilibrium (Extended)
- Optimal design of experiments by combining coarse and fine measurements
- Fluctuations and the limit of predictability in protein evolution
- adabmDCA 2.0 -- a flexible but easy-to-use package for Direct Coupling Analysis
- Impact of phylogeny on the inference of functional sectors from protein sequence data
- Correlated evolution: models and methods
- Small Coupling Expansion for Multiple Sequence Alignment
- Functional bottlenecks can emerge from non-epistatic underlying traits
- adabmDCA: Adaptive Boltzmann machine learning for biological sequences