Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences
arXiv:1512.01801 · doi:10.1093/bioinformatics/btw152
Abstract
Motivation: Single Molecule Real-Time (SMRT) sequencing technology and Oxford Nanopore technologies (ONT) produce reads over 10kbp in length, which have enabled high-quality genome assembly at an affordable cost. However, at present, long reads have an error rate as high as 10-15%. Complex and computationally intensive pipelines are required to assemble such reads. Results: We present a new mapper, minimap, and a de novo assembler, miniasm, for efficiently mapping and assembling SMRT and ONT reads without an error correction stage. They can often assemble a sequencing run of bacterial data into a single contig in a few minutes, and assemble 45-fold C. elegans data in 9 minutes, orders of magnitude faster than the existing pipelines. We also introduce a pairwise read mapping format (PAF) and a graphical fragment assembly format (GFA), and demonstrate the interoperability between ours and current tools. Availability and implementation: https://github.com/lh3/minimap and https://github.com/lh3/miniasm Contact: [email protected]
Identical to the published version except formatting
References in corpus (1)
Cited by in corpus (19)
- Minimap2: pairwise alignment for nucleotide sequences
- Haplotype-resolved de novo assembly with phased assembly graphs
- Nanopore Sequencing Technology and Tools for Genome Assembly: Computational Analysis of the Current State, Bottlenecks and Future Directions
- Technology dictates algorithms: Recent developments in read alignment
- Apollo: A Sequencing-Technology-Independent, Scalable, and Accurate Assembly Polishing Algorithm
- Genome assembly using quantum and quantum-inspired annealing
- BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis
- SeGraM: A Universal Hardware Accelerator for Genomic Sequence-to-Graph and Sequence-to-Sequence Mapping
- diBELLA: Distributed Long Read to Long Read Alignment
- Robust haplotype-resolved assembly of diploid individuals without parental data
- The design and construction of reference pangenome graphs
- SpanSeq: Similarity-based sequence data splitting method for improved development and assessment of deep learning projects
- Linking de novo assembly results with long DNA reads by dnaasm-link application
- RASSA: Resistive Pre-Alignment Accelerator for Approximate DNA Long Read Mapping
- Distributed-Memory Parallel Contig Generation for De Novo Long-Read Genome Assembly
- Robust Seriation and Applications to Cancer Genomics
- Computer Architecture-Aware Optimisation of DNA Analysis Systems
- Accelerating Genome Sequence Analysis via Efficient Hardware/Algorithm Co-Design
- A spectral algorithm for fast de novo layout of uncorrected long nanopore reads