Language Modeling with Reduced Densities
arXiv:2007.03834 · doi:10.32408/compositionality-3-4
Abstract
This work originates from the observation that today's state-of-the-art statistical language models are impressive not only for their performance, but also - and quite crucially - because they are built entirely from correlations in unstructured text data. The latter observation prompts a fundamental question that lies at the heart of this paper: What mathematical structure exists in unstructured text data? We put forth enriched category theory as a natural answer. We show that sequences of symbols from a finite alphabet, such as those found in a corpus of text, form a category enriched over probabilities. We then address a second fundamental question: How can this information be stored and modeled in a way that preserves the categorical structure? We answer this by constructing a functor from our enriched category of text to a particular enriched category of reduced density operators. The latter leverages the Loewner order on positive semidefinite operators, which can further be interpreted as a toy example of entailment.
21 pages; v2: added reference; v3: revised abstract and introduction for clarity; v4: Compositionality version
References in corpus (14)
- The density-matrix renormalization group in the age of matrix product states
- Scaling Laws for Neural Language Models
- MADE: Masked Autoencoder for Distribution Estimation
- Tensor Networks in a Nutshell
- Mathematical Foundations for a Compositional Distributional Model of Meaning
- Quantum Language Model with Entanglement Embedding for Question Answering
- Lectures on Quantum Tensor Networks
- Tight spans, Isbell completions and semi-tropical modules
- Anomaly Detection with Tensor Networks
- Tensor network language model
- Language as a matrix product state
- Entanglement and Tensor Networks for Supervised Image Classification
- A Multi-Scale Tensor Network Architecture for Classification and Regression
- At the Interface of Algebra and Statistics