Fitting Ranked English and Spanish Letter Frequency Distribution in U.S. and Mexican Presidential Speeches
arXiv:1103.2950 · doi:10.1080/09296174.2011.608606
Abstract
The limited range in its abscissa of ranked letter frequency distributions causes multiple functions to fit the observed distribution reasonably well. In order to critically compare various functions, we apply the statistical model selections on ten functions, using the texts of U.S. and Mexican presidential speeches in the last 1-2 centuries. Dispite minor switching of ranking order of certain letters during the temporal evolution for both datasets, the letter usage is generally stable. The best fitting function, judged by either least-square-error or by AIC/BIC model selection, is the Cocho/Beta function. We also use a novel method to discover clusters of letters by their observed-over-expected frequency ratios.
7 figures
References in corpus (4)
Cited by in corpus (9)
- Diminishing Return for Increased Mappability with Longer Sequencing Reads: Implications of the k-mer Distributions in the Human Genome
- Beyond Zipf's Law: The Lavalette Rank Function and its Properties
- Analyses of Baby Name Popularity Distribution in U.S. for the Last 131 Years
- Range-Limited Heaps' Law for Functional DNA Words in the Human Genome
- Approaches to the classification of complex systems: Words, texts, and more
- In silico model of infection of a CD4(+) T-cell by a human immunodeficiency type 1 virus, and a mini-review on its molecular pathophysiology
- Quadratic Term Correction on Heaps' Law
- The 'Letter' Distribution in the Chinese Language
- Beta Rank Function: A Smooth Double-Pareto-Like Distribution