Technology dictates algorithms: Recent developments in read alignment
arXiv:2003.00110 · doi:10.1186/s13059-021-02443-7
Abstract
Massively parallel sequencing techniques have revolutionized biological and medical sciences by providing unprecedented insight into the genomes of humans, animals, and microbes. Modern sequencing platforms generate enormous amounts of genomic data in the form of nucleotide sequences or reads. Aligning reads onto reference genomes enables the identification of individual-specific genetic variants and is an essential step of the majority of genomic analysis pipelines. Aligned reads are essential for answering important biological questions, such as detecting mutations driving various human diseases and complex traits as well as identifying species present in metagenomic samples. The read alignment problem is extremely challenging due to the large size of analyzed datasets and numerous technological limitations of sequencing platforms, and researchers have developed novel bioinformatics algorithms to tackle these difficulties. Importantly, computational algorithms have evolved and diversified in accordance with technological advances, leading to todays diverse array of bioinformatics tools. Our review provides a survey of algorithmic foundations and methodologies across 107 alignment methods published between 1988 and 2020, for both short and long reads. We provide rigorous experimental evaluation of 11 read aligners to demonstrate the effect of these underlying algorithms on speed and efficiency of read aligners. We separately discuss how longer read lengths produce unique advantages and limitations to read alignment techniques. We also discuss how general alignment algorithms have been tailored to the specific needs of various domains in biology, including whole transcriptome, adaptive immune repertoire, and human microbiome studies.
References in corpus (9)
- Minimap2: pairwise alignment for nucleotide sequences
- Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences
- GRIM-Filter: Fast Seed Location Filtering in DNA Read Mapping Using Processing-in-Memory Technologies
- GateKeeper: A New Hardware Architecture for Accelerating Pre-Alignment in DNA Short Read Mapping
- Shouji: A Fast and Efficient Pre-Alignment Filter for Sequence Alignment
- Accelerating Genome Analysis: A Primer on an Ongoing Journey
- Apollo: A Sequencing-Technology-Independent, Scalable, and Accurate Assembly Polishing Algorithm
- SneakySnake: A Fast and Accurate Universal Genome Pre-Alignment Filter for CPUs, GPUs, and FPGAs
- Optimal Seed Solver: Optimizing Seed Selection in Read Mapping
Cited by in corpus (7)
- FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications
- BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis
- SeGraM: A Universal Hardware Accelerator for Genomic Sequence-to-Graph and Sequence-to-Sequence Mapping
- Scrooge: A Fast and Memory-Frugal Genomic Sequence Aligner for CPUs, GPUs, and ASICs
- TargetCall: Eliminating the Wasted Computation in Basecalling via Pre-Basecalling Filtering
- GateKeeper-GPU: Fast and Accurate Pre-Alignment Filtering in Short Read Mapping
- AirLift: A Fast and Comprehensive Technique for Remapping Alignments between Reference Genomes