9 papers
Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation
Vitória Barin Pacela, Shruti Joshi, Isabela Camacho +2
The linear representation hypothesis states that neural network activations encode high-level concepts as linear mixtures. However, under superposition, this encoding is a projecti…
Causality is Key for Interpretability Claims to Generalise
Shruti Joshi, Aaron Mueller, David Klindt +3
Interpretability research on large language models (LLMs) has yielded important insights into model behaviour, yet recurring pitfalls persist: findings that do not generalise, and…
FederatedFactory: Generative One-Shot Learning for Extremely Non-IID Distributed Scenarios
Andrea Moleri, Christian Internò, Ali Raza +4
Federated Learning (FL) enables distributed optimization without compromising data sovereignty. Yet, where local label distributions are mutually exclusive, standard weight aggrega…
The Observer Effect in World Models: Invasive Adaptation Corrupts Latent Physics
Christian Internò, Jumpei Yamaguchi, Loren Amdahl-Culleton +3
Determining whether neural models internalize physical laws as world models, rather than exploiting statistical shortcuts, remains challenging, especially under out-of-distribution…
AI-Generated Video Detection via Perceptual Straightening
Christian Internò, Robert Geirhos, Markus Olhofer +3
The rapid advancement of generative AI enables highly realistic synthetic videos, posing significant challenges for content authentication and raising urgent concerns about misuse.…
Superposition disentanglement of neural representations reveals hidden alignment
André Longon, David Klindt, Meenakshi Khosla
The superposition hypothesis states that single neurons may participate in representing multiple features in order for the neural network to represent more features than it has neu…