5 papers
Emergent retokenization symmetry in large language models: phenomenology and applications
Kanishk Jain, Matthew Day, Tankut Can
Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same…
Semantic Chunking and the Entropy of Natural Language
Weishun Zhong, Doron Sivan, Tankut Can +2
The entropy rate of printed English is famously estimated to be about one bit per character, a benchmark that modern large language models (LLMs) have only recently approached. Thi…
Statistical Mechanics of Semantic Compression
Tankut Can
The basic problem of semantic compression is to minimize the length of a message while preserving its meaning. This differs from classical notions of compression in that the distor…
Random Tree Model of Meaningful Memory
Weishun Zhong, Tankut Can, Antonis Georgiou +3
Traditional studies of memory for meaningful narratives focus on specific stories and their semantic structures but do not address common quantitative features of recall across dif…
Large-scale study of human memory for meaningful narratives
Antonios Georgiou, Tankut Can, Mikhail Katkov +1
The statistical study of human memory requires large-scale experiments, involving many stimuli conditions and test subjects. While this approach has proven to be quite fruitful for…