5 papers
Logit Distance Bounds Representational Similarity
Beatrix M. G. Nielsen, Emanuele Marconato, Luigi Gresele +2
For a broad family of discriminative models that includes autoregressive language models, identifiability results imply that if two models induce the same conditional distributions…
Relational Linear Properties in Language Models: An Empirical Investigation
Giovanni Valer, Luigi Gresele, Marco Bronzini +1
Linear properties are ubiquitous in the representations of language models; however, testing them experimentally remains a challenging task. This work focuses on relational lineari…
When Does Closeness in Distribution Imply Representational Similarity? An Identifiability Perspective
Beatrix M. G. Nielsen, Emanuele Marconato, Andrea Dittadi +1
When and why representations learned by different deep neural networks are similar is an active research topic. We choose to address these questions from the perspective of identif…
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
Emanuele Marconato, Sébastien Lachapelle, Sebastian Weichwald +1
We analyze identifiability as a possible explanation for the ubiquity of linear properties across language models, such as the vector difference between the representations of "eas…
What is causal about causal models and representations?
Frederik Hytting Jørgensen, Luigi Gresele, Sebastian Weichwald
Causal Bayesian networks are 'causal' models since they make predictions about interventional distributions. To connect such causal model predictions to real-world outcomes, we mus…