Simulations, Computations, and Statistics for Longest Common Subsequences
arXiv:1705.06826
Abstract
The length of the longest common subsequences (LCSs) is often used as a similarity measurement to compare two (or more) random words. Below we study its statistical behavior in mean and variance using a Monte-Carlo approach from which we then develop a hypothesis testing method for sequences similarity. Finally, theoretical upper bounds are obtained for the Chvátal-Sankoff constant of multiple sequences.
12 pages, 10 figures