Coding for Composite DNA to Correct Substitutions, Strand Losses, and Deletions
arXiv:2404.12868 · doi:10.1109/ISIT57864.2024.10619202
Abstract
Composite DNA is a recent method to increase the base alphabet size in DNA-based data storage.This paper models synthesizing and sequencing of composite DNA and introduces coding techniques to correct substitutions, losses of entire strands, and symbol deletion errors. Non-asymptotic upper bounds on the size of codes with occurrences of these error types are derived. Explicit constructions are presented which can achieve the bounds.