4 papers · 1 filter
Measuring Optimal Transport in Transformer Depth
Alexandre Quemy
A transformer carries each token's state from layer to layer, and the whole vocabulary carried together forms a cloud that moves with depth. We ask whether a trained network moves…
The Depth Flow of Token Representations Is Nonlinear and Does Not Descend Its Own Density
Alexandre Quemy
A token's representation is carried through the network layer by layer. The whole vocabulary carried together forms a flow. We fit this flow's equation of motion as a discrete Lang…
A Hub of Short Rows Inflates Intrinsic Dimension Estimation of Token Embeddings
Alexandre Quemy
A token-embedding table holds a hub of short rows near its origin, and we show that this cluster biases what nearest-neighbor intrinsic-dimension (ID) estimators report. Because of…
Riemannian Geometry for Pre-trained Language Model Embeddings
Szczepan Konior, Alexandre Quemy, Przemysław Klocek +2
Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in…