Language Trees and Zipping
arXiv:cond-mat/0108530 · doi:10.1103/PhysRevLett.88.048702
Abstract
In this letter we present a very general method to extract information from a generic string of characters, e.g. a text, a DNA sequence or a time series. Based on data-compression techniques, its key point is the computation of a suitable measure of the remoteness of two bodies of knowledge. We present the implementation of the method to linguistic motivated problems, featuring highly accurate results for language recognition, authorship attribution and language classification.
5 pages, RevTeX4, 1 eps figure. In press in Phys. Rev. Lett. (January 2002)
Cited by in corpus (32)
- Colloquium: Criticality and dynamical scaling in living systems
- Universal and accessible entropy estimation using a compression algorithm
- Clustering by compression
- Data compression and learning in time sequences analysis
- Identifying Cover Songs Using Information-Theoretic Measures of Similarity
- Extended Comment on Language Trees and Zipping
- Analysis of the phase transition in the Ising ferromagnet using a Lempel-Ziv string parsing scheme and black-box data-compression utilities
- Equilibrium (Zipf) and Dynamic (Grasseberg-Procaccia) method based analyses of human texts. A comparison of natural (english) and artificial (esperanto) languages
- Entropy and hierarchical clustering: characterising the morphology of the urban fabric in different spatial cultures
- Authorship Analysis based on Data Compression
- Compression ratios based on the Universal Similarity Metric still yield protein distances far from CATH distances
- Measuring complexity with zippers
- Vicsek Model by Time-Interlaced Compression: a Dynamical Computable Information Density
- Hierarchy of Scales in Language Dynamics
- Complexity of Networks
- Alignment-free comparison of next-generation sequencing data using compression-based distance measures
- Artificial Sequences and Complexity Measures
- The similarity metric
- Efficiency of the Moscow Stock Exchange before 2022
- Dictionary based methods for information extraction
- Normalized Information Distance
- From form to information: Analysing built environments in different spatial cultures
- Quantifying spatio-temporal patterns in classical and quantum systems out of equilibrium
- Quantifying Local Randomness in Human DNA and RNA Sequences Using Erdos Motifs
- On the Ziv-Merhav theorem beyond Markovianity
- Information Distance in Multiples
- Diversity, competition, extinction: the ecophysics of language change
- God (), the first small world network
- Sublinear Algorithms for Approximating String Compressibility
- Ziv-Merhav estimation for hidden-Markov processes
- Signal processing and statistical methods in analysis of text and DNA
- Relative entropy via non-sequential recursive pair substitutions