Minimap2: pairwise alignment for nucleotide sequences
arXiv:1708.01492 · doi:10.1093/bioinformatics/bty191
Abstract
Motivation: Recent advances in sequencing technologies promise ultra-long reads of 100 kilo bases (kb) in average, full-length mRNA or cDNA reads in high throughput and genomic contigs over 100 mega bases (Mb) in length. Existing alignment programs are unable or inefficient to process such data at scale, which presses for the development of new alignment algorithms. Results: Minimap2 is a general-purpose alignment program to map DNA or long mRNA sequences against a large reference database. It works with accurate short reads of 100bp in length, 1kb genomic reads at error rate 15%, full-length noisy Direct RNA or cDNA reads, and assembly contigs or closely related full chromosomes of hundreds of megabases in length. Minimap2 does split-read alignment, employs concave gap cost for long insertions and deletions (INDELs) and introduces new heuristics to reduce spurious alignments. It is 3-4 times faster than mainstream short-read mappers at comparable accuracy and 30 times faster at higher accuracy for both genomic and mRNA reads, surpassing most aligners specialized in one type of alignment. Availability and implementation: https://github.com/lh3/minimap2 Contact: [email protected]
The final submitted version
References in corpus (1)
Cited by in corpus (40)
- Haplotype-resolved de novo assembly with phased assembly graphs
- CoverM: Read alignment statistics for metagenomics
- Technology dictates algorithms: Recent developments in read alignment
- Fast Characterization of Segmental Duplications in Genome Assemblies
- Accelerating Genome Analysis: A Primer on an Ongoing Journey
- FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications
- Apollo: A Sequencing-Technology-Independent, Scalable, and Accurate Assembly Polishing Algorithm
- Genome assembly using quantum and quantum-inspired annealing
- ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads
- SneakySnake: A Fast and Accurate Universal Genome Pre-Alignment Filter for CPUs, GPUs, and FPGAs
- RawHash: Enabling Fast and Accurate Real-Time Analysis of Raw Nanopore Signals for Large Genomes
- TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
- BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis
- Internal Pattern Matching Queries in a Text and Applications
- BWT construction and search at the terabase scale
- The Parallelism Motifs of Genomic Data Analysis
- RawHash2: Mapping Raw Nanopore Signals Using Hash-Based Seeding and Adaptive Quantization
- Scrooge: A Fast and Memory-Frugal Genomic Sequence Aligner for CPUs, GPUs, and ASICs
- diBELLA: Distributed Long Read to Long Read Alignment
- The design and construction of reference pangenome graphs
- Minimum error correction-based haplotype assembly: considerations for long read data
- TargetCall: Eliminating the Wasted Computation in Basecalling via Pre-Basecalling Filtering
- wgatools: an ultrafast toolkit for manipulating whole genome alignments
- New strategies to improve minimap2 alignment accuracy
- Genome assembly in the telomere-to-telomere era
- RAPIDx: High-performance ReRAM Processing in-Memory Accelerator for Sequence Alignment
- CiMBA: Accelerating Genome Sequencing through On-Device Basecalling via Compute-in-Memory
- GateKeeper-GPU: Fast and Accurate Pre-Alignment Filtering in Short Read Mapping
- RASSA: Resistive Pre-Alignment Accelerator for Approximate DNA Long Read Mapping
- Numeric Lyndon-based feature embedding of sequencing reads for machine learning approaches
- NMP-PaK: Near-Memory Processing Acceleration of Scalable De Novo Genome Assembly
- MetaCompass: Reference-guided Assembly of Metagenomes
- A mixture model for determining SARS-Cov-2 variant composition in pooled samples
- Reconstructing Latent Orderings by Spectral Clustering
- ChloroScan: Recovering plastid genome bins from metagenomic data
- A mapping-free NLP-based technique for sequence search in Nanopore long-reads
- DNA Pre-alignment Filter using Processing Near Racetrack Memory
- Computer Architecture-Aware Optimisation of DNA Analysis Systems
- PASS: De novo assembler for short peptide sequences
- Accelerating Genome Sequence Analysis via Efficient Hardware/Algorithm Co-Design