most citedSymmetry in language statistics shapes the geometry of model representations

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.LG20261 cited

Symmetry in language statistics shapes the geometry of model representations

Dhruva Karkada, Daniel J. Korchinski, Andres Nava +2

The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth on…

cs.LG2026

Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?

Marta Aparicio Rodriguez, Anastasia Borovykh, Grigorios A. Pavliotis +1

Generative models have a persistent limitation: their tendency to memorize training data can create legal liabilities and erode creative diversity. Understanding which samples are…

cs.LG2026

Learn from your own latents and not from tokens: A sample-complexity theory

Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart

Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological lea…

cs.LG2026

Sampling Data with Chains of Forward-Backward Diffusion Steps

Hyunmo Kang, Noam Itzhak Levi, Corinna Elena Wegner +2

Sampling from learned high-dimensional distributions is a foundational computational problem. We introduce U-turn chains: Markov chains obtained by iterating short forward-backward…

cs.CL2025

On the Emergence of Linear Analogies in Word Embeddings

Daniel J. Korchinski, Dhruva Karkada, Yasaman Bahri +1

Models such as Word2Vec and GloVe construct word embeddings based on the co-occurrence probability of words and in text corpora. The resulting vectors not on…