New stopping criteria for segmenting DNA sequences
arXiv:physics/0104026 · doi:10.1103/PhysRevLett.86.5815
Abstract
We propose a solution on the stopping criterion in segmenting inhomogeneous DNA sequences with complex statistical patterns. This new stopping criterion is based on Bayesian Information Criterion (BIC) in the model selection framework. When this stopping criterion is applied to a left telomere sequence of yeast Saccharomyces cerevisiae and the complete genome sequence of bacterium Escherichia coli, borders of biologically meaningful units were identified (e.g. subtelomeric units, replication origin, and replication terminus), and a more reasonable number of domains was obtained. We also introduce a measure called segmentation strength which can be used to control the delineation of large domains. The relationship between the average domain size and the threshold of segmentation strength is determined for several genome sequences.
4 pages, 4 figures, Physical Review Letters, to appear
Cited by in corpus (11)
- Will the US Economy Recover in 2010? A Minimal Spanning Tree Study
- Stable Distributions in Stochastic Fragmentation
- Beyond Zipf's Law: The Lavalette Rank Function and its Properties
- Phase Transition in a Random Fragmentation Problem with Applications to Computer Science
- Simplifying the mosaic description of DNA sequences
- Bridging stylized facts in finance and data non-stationarities
- Non-parametric segmentation of non-stationary time series
- Extending the Recursive Jensen-Shannon Segmentation of Biological Sequences
- Causal Links Between US Economic Sectors
- DNA Segmentation as A Model Selection Process
- Macroeconomic Phase Transitions Detected from the Dow Jones Industrial Average Time Series