Inferring processes underlying B-cell repertoire diversity
arXiv:1502.03136 · doi:10.1098/rstb.2014.0243
Abstract
We quantify the VDJ recombination and somatic hypermutation processes in human B-cells using probabilistic inference methods on high-throughput DNA sequence repertoires of human B-cell receptor heavy chains. Our analysis captures the statistical properties of the naive repertoire, first after its initial generation via VDJ recombination and then after selection for functionality. We also infer statistical properties of the somatic hypermutation machinery (exclusive of subsequent effects of selection). Our main results are the following: the B-cell repertoire is substantially more diverse than T-cell repertoires, due to longer junctional insertions; sequences that pass initial selection are distinguished by having a higher probability of being generated in a VDJ recombination event; somatic hypermutations have a non-uniform distribution along the V gene that is well explained by an independent site model for the sequence context around the hypermutation site.
acknowledgement added
References in corpus (2)
Cited by in corpus (26)
- IGoR: a tool for high-throughput immune repertoire analysis
- OLGA: fast computation of generation probabilities of B- and T-cell receptor amino acid sequences and motifs
- Computational strategies for dissecting the high-dimensional complexity of adaptive immune repertoires
- Measuring the sequence-affinity landscape of antibodies with massively parallel titration curves
- Consistency of VDJ rearrangement and substitution parameters enables accurate B cell receptor sequence annotation
- Predicting the spectrum of TCR repertoire sharing with a data-driven model of recombination
- Likelihood-based inference of B-cell clonal families
- Genesis of the alpha beta T-cell receptor
- Deep generative selection models of T and B cell receptor repertoires with soNNia
- Per-sample immunoglobulin germline inference from B cell receptor deep sequencing data
- Augmenting adaptive immunity: progress and challenges in the quantitative engineering and analysis of adaptive immune receptor repertoires
- Quantitative Immunology for Physicists
- Linguistically inspired roadmap for building biologically reliable protein language models
- Tuning environmental timescales to evolve and maintain generalists
- Learning the heterogeneous hypermutation landscape of immunoglobulins from high-throughput repertoire data
- Mouse T cell repertoires as statistical ensembles: overall characterization and age dependence
- repgenHMM: a dynamic programming tool to infer the rules of immune receptor generation from sequence data
- A Bayesian Phylogenetic Hidden Markov Model for B Cell Receptor Sequence Analysis
- Learning the statistics and landscape of somatic mutation-induced insertions and deletions in antibodies
- Calculating Germinal Centre Reactions
- Mathematical Characterization of Private and Public Immune Repertoire Sequences
- ImmunoLingo: Linguistics-based formalization of the antibody language
- Survival analysis of DNA mutation motifs with penalized proportional hazards
- Statistical mechanics of clonal expansion in lymphocyte networks modelled with slow and fast variables
- Branching Random Walks on Binary Strings for Evolutionary Processes
- Fierce selection and interference in B-cell repertoire response to chronic HIV-1