4 papers
Token embeddings violate the manifold hypothesis
Michael Robinson, Sourya Dey, Tony Chiang
A full understanding of the behavior of a large language model (LLM) requires our grasp of its input token space. If this space differs from our assumptions, our comprehension of a…
A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census
John M. Abowd, Tamara Adams, Robert Ashmead +12
We show that individual, confidential microdata records from the 2010 U.S. Census of Population and Housing can be accurately reconstructed from the published tabular summaries. Ni…
LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding
Amirhossein Kazerouni, Soroush Mehraban, Michael Brudno +1
Implicit Neural Representations (INRs) are proving to be a powerful paradigm in unifying task modeling across diverse data domains, offering key advantages such as memory efficienc…
The structure of the token space for large language models
Michael Robinson, Sourya Dey, Shauna Sweet
Large language models encode the correlational structure present in natural language by fitting segments of utterances (tokens) into a high dimensional ambient latent space upon wh…