3 papers
cs.CL2025
Token embeddings violate the manifold hypothesis
Michael Robinson, Sourya Dey, Tony Chiang
A full understanding of the behavior of a large language model (LLM) requires our grasp of its input token space. If this space differs from our assumptions, our comprehension of a…
cs.LG2025
LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding
Amirhossein Kazerouni, Soroush Mehraban, Michael Brudno +1
Implicit Neural Representations (INRs) are proving to be a powerful paradigm in unifying task modeling across diverse data domains, offering key advantages such as memory efficienc…
math.DG2024
The structure of the token space for large language models
Michael Robinson, Sourya Dey, Shauna Sweet
Large language models encode the correlational structure present in natural language by fitting segments of utterances (tokens) into a high dimensional ambient latent space upon wh…