activity
20242026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation

Vitória Barin Pacela, Shruti Joshi, Isabela Camacho +2

The linear representation hypothesis states that neural network activations encode high-level concepts as linear mixtures. However, under superposition, this encoding is a projecti…

cs.LG2026

Causality is Key for Interpretability Claims to Generalise

Shruti Joshi, Aaron Mueller, David Klindt +3

Interpretability research on large language models (LLMs) has yielded important insights into model behaviour, yet recurring pitfalls persist: findings that do not generalise, and…

cs.LG2026

FederatedFactory: Generative One-Shot Learning for Extremely Non-IID Distributed Scenarios

Andrea Moleri, Christian Internò, Ali Raza +4

Federated Learning (FL) enables distributed optimization without compromising data sovereignty. Yet, where local label distributions are mutually exclusive, standard weight aggrega…

cs.LG2026

The Observer Effect in World Models: Invasive Adaptation Corrupts Latent Physics

Christian Internò, Jumpei Yamaguchi, Loren Amdahl-Culleton +3

Determining whether neural models internalize physical laws as world models, rather than exploiting statistical shortcuts, remains challenging, especially under out-of-distribution…

cs.LG2025

Superposition disentanglement of neural representations reveals hidden alignment

André Longon, David Klindt, Meenakshi Khosla

The superposition hypothesis states that single neurons may participate in representing multiple features in order for the neural network to represent more features than it has neu…

cs.LG2025

Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research

Patrik Reizinger, Randall Balestriero, David Klindt +1

Self-Supervised Learning (SSL) powers many current AI systems. As research interest and investment grow, the SSL design space continues to expand. The Platonic view of SSL, followi…